The Fault-to-Fix Execution Gap
Detection is digitized. Records are digitized. The work between the alarm and the completed fix is not, and that is where availability is lost.
Working on a live operating problem? See AIMMS Do the Work
Critical infrastructure has digitized two things very well over the past three decades. It digitized detection: SCADA systems, historians, condition monitoring, and alarm management now flag a fault within seconds of it happening. And it digitized records: CMMS and EAM platforms hold the work orders, asset registers, and maintenance history that auditors and reliability engineers depend on.
What sits between those two achievements is the repair itself, and almost none of it is digitized. The fault is detected by software. The fix is recorded in software. Everything in the middle runs on phone calls, shared-drive searches, tribal memory, and the patience of whoever is on call. That middle is where availability is actually lost.
The execution gap is the work required to turn scattered evidence into a repair that can move
Figure 1. The execution gap is the work required to turn scattered evidence into a repair that can move. Eight separate sources converge on one question: what is blocking this repair now. The sources are alarm and event context, historian trend, asset identity, prior work, current manual revision, safety requirements, inventory truth, and supplier and approval state. From that question the work advances through fault, understand, prepare, repair, return, and learn.
Defining the execution gap
The fault-to-fix execution gap is the elapsed time, and the lost availability, between two well-instrumented events: the moment a system detects a fault, and the moment the asset returns to service with a complete, verified record of what was done. The endpoints are automated. The span between them is coordinated by hand, usually by the same technician who is supposed to be fixing the machine.
Three clarifications make the definition useful.
First, it is not a detection problem. Most operations already know about failures fast. Adding more sensors shortens the time to the alarm, not the time from the alarm to a running asset.
Second, it is not a record-keeping problem. The CMMS will eventually hold a closed work order. The gap is everything the work order does not capture: the hunting, the waiting, and the chasing that happened before anyone could type “complete.”
Third, it is not primarily a skills problem. Good technicians diagnose and repair well. The gap consumes their time with work that is not diagnosis and is not repair: retrieving information, confirming parts, waiting on suppliers, and waiting on signatures.
A pump trips at 2 AM
Consider a concrete case. A critical pump at a treatment plant trips at 2:07 AM. SCADA raises the alarm within seconds. The operator acknowledges it and calls the on-call technician. So far, the digitized part of the process has worked exactly as designed.
The technician arrives at 2:40 with a fault code and not much else. The first task is not repair. It is reconstruction. What did the discharge pressure and motor current look like in the minutes before the trip? That lives in the historian, which the technician may not have access to from the floor. Has this pump tripped this way before? That answer is split between terse CMMS closeout notes and the memory of a senior colleague who is asleep.
Next comes the manual. The OEM troubleshooting guide is a PDF somewhere on a shared drive, or a binder in a cabinet, or neither. The technician narrows the likely cause to a failing mechanical seal, partly from experience and partly from a call that wakes the senior colleague at 3:15.
Now the part. “Mechanical seal” is not enough to order one. The technician needs the exact seal for this pump model, this shaft size, this service. The nameplate is corroded. The CMMS bill of materials was never updated after a previous overhaul substituted a different seal. Identifying the correct part number takes another hour of cross-referencing drawings and old purchase records.
The CMMS shows two seals in stock. The storeroom shelf holds one, and it is the wrong variant. The count drifted months ago and nobody caught it, because nobody needed that seal until tonight.
So the repair now waits on the outside world. At 8:00 AM someone starts calling suppliers for stock and lead time. Quotes arrive by email over the next few hours. The purchase exceeds the technician’s spend authority, so it waits again for a manager who is in meetings until noon. The approval takes thirty seconds to grant and five hours to obtain.
The seal arrives the next morning. The physical replacement takes under three hours: isolate, lock out, pull the pump, swap the seal, reassemble, test, return to service. At the end of a long shift, the technician types a few lines into the work order. The next person to face this failure will inherit almost none of what was learned tonight.
Total elapsed time: roughly thirty hours. Time with tools on the asset: perhaps three.
Where the hours actually go
That ratio is the execution gap made visible. Break the elapsed time into its components and a pattern emerges that most maintenance leaders will recognize from their own operations.
- Diagnosis. Working from a fault code to a probable cause. Fast when the failure mode is familiar, slow when it is not, and heavily dependent on who happens to be on call.
- Information search. Finding trends, drawings, manuals, and prior work orders scattered across the historian, the CMMS, shared drives, and filing cabinets. This is retrieval, not judgment, yet it consumes skilled hours.
- Expert assistance. Phone calls to the one person who has seen this failure before. The knowledge exists in the organization. It is just not written down anywhere the technician can reach at 3 AM.
- Part identification. Translating “this component failed” into an exact, orderable part number, against stale bills of materials and undocumented substitutions.
- Inventory confirmation. Discovering whether the system count matches the shelf. Every phantom stock record converts a same-day repair into a multi-day procurement.
- Supplier response. Requests for stock, price, and lead time that move at the speed of business hours, voicemail, and inbox queues.
- Approval. Spend and work authorizations that sit in someone’s queue. The decision is usually easy. The routing is what takes hours.
- Documentation. Closeout done from memory at the end of a shift, thin enough that it rarely helps the next repair.
Wrench time, the portion where a skilled person is physically restoring the asset, is a minority of the total in most emergency repairs. The rest is coordination. Availability is lost not because repairs are slow, but because everything around the repair is.
Why the gap persists
If the gap is this visible in hindsight, why has it survived decades of software investment? Three reasons.
Fragmented systems
The historian, the CMMS, the parts catalog, the procurement system, the approval workflow, and the supplier’s inbox each do their own job adequately. None of them carries the repair across the others. The technician is the integration layer, walking data from one system to the next by hand. Every boundary between systems becomes a place where the repair can stall.
Tribal knowledge
The most valuable diagnostic information in most plants is not in any system. It is in the heads of senior technicians: which failure modes this specific asset favors, which substitute part actually fits, which supplier answers the phone at night. When those people retire or move on, the organization repeats old diagnoses at full price. Closeout notes written in a hurry do not capture what they knew.
Coordination is nobody’s job description
Planners plan next week’s scheduled work. Operators run the process. Buyers process requisitions in order. Managers approve what reaches them. Emergency repair coordination, the urgent threading of evidence, parts, suppliers, and approvals through a single incident, belongs to no role. So it defaults to the technician, who does it serially, one phone call at a time, while the asset sits idle. Because each incident is treated as a one-off scramble, nothing about the scramble ever improves.
What closing the gap requires
None of this is fixed by another dashboard or another alarm. Closing the execution gap means treating the span between fault and documented fix as a single piece of work, and giving it what any well-run work gets: information, sequence, visibility, and follow-through.
- Evidence brought to the point of work. The trends before the trip, the asset’s failure history, the relevant manual sections, and the record of past fixes should be assembled and waiting when the technician arrives, not hunted down at 3 AM.
- Guided next steps. Troubleshooting should draw on documented history and known failure modes for that specific asset, so the quality of the diagnosis depends less on who happens to be on call.
- Blocker visibility. At any moment, everyone involved should be able to see exactly what the repair is waiting on: a part, a stock confirmation, a quote, an approval. What is visible gets expedited. What is invisible just waits.
- Digital work moving in parallel with physical work. Verifying the exact part, confirming real stock, contacting suppliers, gathering quotes, and routing approvals are digital tasks. They should proceed while the technician isolates and disassembles the asset, within permissions the operator explicitly defines, instead of forming a serial queue afterward.
- Closeout that feeds the next repair. Verified documentation of cause, parts, and procedure should make the next occurrence of this failure faster. That is how the gap shrinks over time instead of resetting with every incident.
This is the work EQUA AIMMS is built to do: an autonomous operating layer for the digital side of the repair, connected read-only to SCADA and historian data, with no control-system writes and autonomy boundaries set by the customer. Technicians repair the asset. The system carries the evidence, the blockers, the coordination, and the record.
Whatever tools you use, the starting point is the same. Take your last ten emergency repairs and separate wrench time from elapsed time. The difference is your execution gap, and it is almost certainly the largest source of recoverable availability you are not currently managing.