← Back to Resources

Guide · July 23, 2026 · 9 min read

SCADA Saw the Alarm. What Happens Next?

Detection is a solved problem. The alarm is accurate, timely, and useless on its own. This is a guide to the eight separate places a water utility has to reach after an alarm, and why assembling them is still a manual job.

Srikant Naidu, Founder, EQUA AI · Updated August 13, 2026

Working on a live operating problem? See AIMMS Do the Work

eight distinct evidence sources converging through separate paths on one central alarm
After the alarm · Eight evidence sources have to converge

The alarm did its job.

A pump vibration crossed its limit at 2:13 a.m. SCADA flagged it in under a second. The historian captured the change. The operator on shift knows something is wrong, knows which asset, and knows when.

Now the harder work starts.

What caused it? Has it happened before? Can the pump keep running until morning? What should the technician inspect first? Which procedure applies to this machine, at this revision? Is the part we are probably going to need actually on the shelf, or only in the system? Is the standby available? Who has to approve the spend if we have to buy it tonight?

Not one of those questions is answered by the alarm. Every one of them has to be answered before anybody touches the equipment. And each answer lives in a different system, owned by a different function, reached in a different way.

This guide is about that interval. It is the least instrumented part of the maintenance process and, in most utilities, the largest.

SCADA is not the problem, and this is not a criticism of it

It is worth being clear about this early, because the argument gets misread as an attack on control systems and it is the opposite.

SCADA is excellent at what it does. It monitors process variables continuously, it controls equipment reliably, it raises alarms accurately, and it does all of that with the determinism a physical process requires. A historian is excellent at retaining operating data at resolution. A computerised maintenance management system is good at recording and organising maintenance work, which is exactly what a system of record should be. Manuals define the equipment. Inventory knows what should be in stock.

Each of these tools does its own job well. None of them was designed to do the job that sits between them.

That job has a name in most plants and it is a person’s name.

The evidence that this is where the loss is

The pattern shows up clearly in sector data. Black and Veatch’s 2026 Water Report, surveying more than 600 United States water sector stakeholders, found that 59 percent of respondents have a data or digital strategy and 70 percent say they collect enough data, but only 19 percent say they leverage it effectively.

That gap is usually read as an analytics problem, and analytics vendors read it that way with enthusiasm. But look at what the same respondents say about the specific data a repair depends on. Only 7 percent rate the asset information in their enterprise asset management or maintenance management system as very good. Only 8 percent say the same of asset operations and maintenance costs, and only 10 percent of asset condition and performance.

Those are not analytics inputs. They are the fields somebody opens at 2 a.m. And in a majority of utilities they are rated average at best, which is a polite way of saying they will not be trusted without verification, which is a longer way of saying somebody will go and check.

There is a third figure worth putting beside those. Asked where simulation modelling offers the most value, 68 percent of respondents pointed to asset management and moving from schedule-based to preventative and predictive maintenance. The sector knows where the value is. What it lacks is not ambition or data. It is the connective step.

The eight inputs

Here is what actually has to converge before a technician can do useful work on that 2:13 a.m. pump. Not conceptually. Literally, on a specific night, in a specific plant.

FIG. 1

Eight separate sources have to converge before anyone can act on one alarm

Figure 1. Eight separate sources have to converge before anyone can act on one alarm. Eight input boxes arranged at uneven left offsets, each with a curved connector leading to a single junction marked with a question mark. The inputs are: the SCADA event itself, the historian trend before it, the asset work history, the manual at the correct revision, the previous repair on this asset, technician evidence from the field, the part identity and real stock, and supplier terms and lead time. The junction feeds one question: what do we do next. Below, a rail of five stages runs left to right: understand, prepare, repair, return, learn.

The uneven starting positions are deliberate. These sources are not arranged in a row waiting to be read. They are in eight systems with eight different access paths, and the convergence is performed by a person, usually the most experienced one available, usually at night.

Take them one at a time, because the specific friction in each is different and the remedies are different.

1. The SCADA event

The easiest one, and still not free. The alarm names a tag. The tag may not obviously map to the asset as the maintenance system knows it. The same physical pump is often P-101 in the maintenance system, a different tag in SCADA, a third string on the nameplate, and “the east influent pump” in a decade of technician notes. Before any history can be assembled, somebody reconciles those identities by hand, and every search that misses an alias silently misses part of the story.

2. The historian trend before the event

Frequently the most diagnostic single input, and frequently the least accessible. The trend for the hour before the trip distinguishes a gradual thermal rise from an instantaneous electrical event, which points at completely different causes. It is also behind a client application and a login the on-call technician may not have, at 2 a.m., from a truck.

3. The asset work history

What has been done to this machine. In principle a solved problem. In practice, the last repair is a two-line entry that says “replaced seal”, which tells you the outcome and nothing about the reasoning, and the entries before that are spread across a rename, a system migration, and two spellings of the same asset.

4. The manual, at the correct revision

Not the manual for this model. The manual matching what is actually installed after the 2021 rebuild with a substitute impeller. Those are different documents and the difference is discovered by fitting the wrong thing.

5. The previous repair on this asset

Distinct from work history, and this is the distinction most systems miss. History is a list of events. The previous repair is a narrative: what was suspected, what was checked, what was eliminated, what worked. Almost no utility holds this, because no system asks for it at a moment when anyone would answer.

6. Technician evidence from the field

What the person at the asset is perceiving right now. The photograph of the coupling. The reading from the handheld. The observation that it smells hot. This is generated at the asset and, in most plants, transmitted by phone call to somebody who writes part of it down.

7. Part identity and real stock

Two different facts routinely treated as one. Which part fits this asset in its current configuration, and whether a usable one is physically on the shelf. Inventory answers a question adjacent to both: what the system believes it has. The gap between the belief and the shelf is a familiar and expensive discovery.

8. Supplier terms and lead time

Only relevant when the part is not there, at which point it is the only thing that matters. It lives in a purchasing system, an email thread, and a contract nobody has open.

Who does the converging

In every plant, the answer is a person. Usually the most experienced person on shift, because the convergence requires knowing where things are, which is itself the scarce knowledge.

That has three consequences worth naming.

It is the wrong use of the scarcest resource. The most valuable diagnostic capability in the building is spending the first ninety minutes on retrieval rather than diagnosis.

It is invisible in every metric. Time-to-repair measures the interval. Nothing decomposes it. So the retrieval cost never appears in a budget conversation, and the improvement that would help most is never funded because it was never measured.

It degrades exactly when it is needed most. The convergence is hardest at 2 a.m., on an unfamiliar asset, during a storm event, with the experienced person unavailable. Which is to say, during the events that actually matter.

What good looks like

A useful standard, and one you can test against what you have now.

The evidence arrives with the fault. Nobody should have to know that a relevant record exists in order to receive it. If retrieval depends on knowing what to search for, retrieval is gated on the knowledge that is scarcest.

Every line names its source and its age. A working context that presents a value without saying it came from the historian and is four hours old is worse than no context, because it invites a decision it cannot support.

Absence is shown, not filled. If the manual revision is unknown or a source is unreachable, that is the most important thing on the screen. A system that quietly substitutes a plausible value for a missing one has produced confidence without information, which is the one failure mode a maintenance team cannot absorb.

Diagnosis stays with people. Ranked likely causes are a starting point for a qualified person to accept or reject. Anything more assertive is asking software to hold a responsibility it cannot carry.

Preparation runs in parallel with diagnosis. Part identification, stock verification, and the approval route do not need to wait for the cause to be confirmed. Much of the waiting in a typical repair is sequential only by habit.

The distinction that matters: understanding versus preparing

Most discussion of AI in maintenance stops at the first of these. Diagnosis is the interesting problem, so it gets the attention.

But diagnosis is not what keeps assets down. Consider the 2:13 a.m. pump with a confirmed diagnosis at 4 a.m.: the mechanical seal has failed. Excellent. The asset is still down, and now the questions are which seal kit matches the installed configuration, whether one is on the shelf, whether the count is right, who can authorise the purchase if it is not, when the supplier can deliver, when the crew and the isolation and the process window agree, and who signs the return to service.

None of those is a diagnostic question. All of them are coordination. And in a public utility, several of them are governed by procurement rules that exist for good reasons and add days.

The alarm-to-action gap is mostly not an intelligence problem. It is a coordination problem wearing an intelligence problem’s clothes.

Where AIMMS fits

EQUA AIMMS is built for the interval this guide describes. It starts where the alarm stops: assembling the permitted evidence around a fault into one working context with each line traced to its source and its age, supporting the team’s understanding of what is happening, and then moving the digital work around the repair, in parts, sourcing, approvals and records, under the customer’s own authority rules.

Three boundaries define it as much as the capability does. Context from SCADA and the historian is read only, and AIMMS holds no write path to a PLC, DCS or SCADA system and no ability to issue setpoints, restarts, interlocks or actuator commands. Qualified people diagnose, repair, and authorise return to service. And your existing systems of record stay authoritative, receiving only what you have configured, tested and approved.

The question to take to your own plant

Pick your last significant unplanned failure. Reconstruct the timeline honestly, and mark the moment the technician had everything needed to start work.

Then measure how much of the total downtime happened before that moment.

In most plants the answer is uncomfortable, and it is also the most actionable number available, because it is the part nobody is working on.

Sources

  • Black and Veatch, 2026 Water Report, 15th annual edition, survey of more than 600 United States water sector stakeholders, June 2026. Figures cited: 59 percent with a data or digital strategy, 70 percent collecting enough data, 19 percent leveraging it effectively; asset information rated very good by 7 percent for enterprise asset management and maintenance management records, 8 percent for asset operations and maintenance costs, 10 percent for asset condition and performance; 68 percent identifying asset management and the move from schedule-based to preventative and predictive maintenance as where simulation modelling offers the most value.

Turn this idea into a facility-specific decision.

Bring one recurring failure or stuck workflow. The path is deliberately focused:

  1. 01

    Intake

    Complete a short qualification intake.

  2. 02

    Working session

    Map the delay and control boundary in 20 minutes.

  3. 03

    First-scope decision

    Decide whether a credible facility-specific first scope exists.

See AIMMS Do the Work