← Back to Resources

Guide · September 7, 2026 · 18 min read

The Maintenance Benchmark Provenance Ledger: Published, Practitioner Opinion, or Folklore

Figure by figure: where each maintenance benchmark actually comes from, what was measured, the sample and the year, and whether it traces to a primary source at all. EQUA's own four numbers are graded by the same rule.

Srikant Naidu, Founder, EQUA AI

Working on a live operating problem? Book Your 20-Minute Assessment

a maintenance office desk at night under a single lamp, an open handwritten equipment logbook ruled into columns, stacked ring binders of history, loose printed tables held by a steel rule and a pair of reading glasses, filing cabinets in shadow behind
Equipment history - where a number has to be traceable back to

Several of the most-quoted maintenance benchmarks are not measurements at all, and several that are get quoted in a form their source does not support. The US Department of Energy’s Operations and Maintenance Best Practices Guide Release 3.0 (August 2010) really does publish 10 times return, 70 to 75 percent elimination of breakdowns and 35 to 45 percent reduction in downtime for predictive maintenance, but attributes them only to unnamed “independent surveys”. The work-sampling paper published in the Journal of Industrial Intelligence in 2024 measured 28 percent wrench time at one chemical plant in the first half of 2023, not the 55 percent quoted as world class.

Key takeaways

  • The four-part test. Four things must be readable in the primary before a figure ships: the number, the year, the population it was drawn from, and the exact question it answered. Most circulating benchmarks fail on the last two.
  • The DOE predictive-maintenance figures exist. They are not invented. The criticism that lands is that a 2010 federal guide sourced them to unidentified surveys, and the citation chain since has added no measurement.
  • Wrench-time benchmarks disagree because the definitions disagree. Doc Palmer and Torbjörn Idhammar, two of the most-cited practitioners on the metric, both warn against using it the way it is normally used.
  • Press releases broaden scoped results. The Black and Veatch 2026 Water Report press release attaches 66 percent to aging infrastructure. In the report, that number answers a narrower question.
  • Denominators travel with the number. Uptime Institute’s 40 percent is the share of organizations that had a human-error outage in three years, not the share of outages caused by human error.
  • Publication year is not measurement year. The 28 percent wrench-time paper appeared in 2024. The observations were taken in the first half of 2023.
  • Practitioner benchmark is a category, not an insult. 90 percent PM compliance and one planner per 20 technicians are useful conventions. They are not findings.
  • We graded our own four numbers by the same rule, including the clock the 96 percent figure runs on.

How this ledger grades a figure

Three verdicts, and only three.

Published and traceable: a named document, from a named publisher, with a date, states the number and makes the population and the question legible. A sceptical engineer can check it in a minute.

Practitioner benchmark, not a study: a widely agreed operating target published by reliability bodies, planning authors and consultancies. It encodes real experience, but it is not a measurement across a defined population, and no such measurement appears when you follow the citation chain back.

Cannot be traced to a primary: the figure carries an authoritative-sounding attribution the attributed body does not publish, or the chain terminates in another secondary source.

One further pattern needs naming: a figure that is genuinely traceable, then quoted in a form the primary does not support. That is a scope error, and it is more dangerous than folklore, because the citation survives checking while the claim does not.

The ledger

Figure as usually quotedWhere it actually comes fromWhat was measuredSample and yearVerdict
“World-class wrench time is 55%”No study. A best-practice target from planning authors, notably Doc Palmer, repeated by vendors and consultanciesNothing. A target, not an observed distributionNone statedPractitioner benchmark, not a study
“Typical wrench time is 25% to 35%”Attributed most often to Smith and Mobley, Rules of Thumb for Maintenance and Reliability Engineers (2007), a paid text. Idhammar states the same band as what a wrench-time study “will typically show”A rule of thumb from consulting experienceNone statedPractitioner benchmark, not a study
“Wrench time is 28%”Rahman, Journal of Industrial Intelligence 2(3), pp. 172 to 188, online 30 September 2024Work sampling: share of random observations in which a technician was working directly on the job320 sampled maintenance activities over three months at one chemical processing plant, 50 maintenance personnel across four crews and eight trade classifications. Observed in the first half of 2023, published 2024Published and traceable, single site
“10x return on predictive maintenance”, “70% to 75% elimination of breakdowns”, “35% to 45% reduction in downtime”US DOE Federal Energy Management Program, Operations and Maintenance Best Practices Guide Release 3.0, section 5.4, page 5.4, August 2010“Industrial average savings” indicated by “independent surveys” that are not named and are not in the chapter’s referencesNot stated. Published 2010Published and traceable to the guide; the guide’s own evidence is not
“90% PM compliance is world class”Reliability bodies, consultancies and CMMS vendor documentationA completion target against a compliance window, defined differently by each publisherNone statedPractitioner benchmark, not a study
“A planner for every 20 technicians”Doc Palmer, Plant Services, 25 June 2024. Stated form: “one planner for every 20-30 craftspersons”A ratio argued from planning economics, not sampled from plantsNone statedPractitioner benchmark, not a study
“Sixty-six percent of respondents cite aging infrastructure as a top challenge”Black and Veatch 2026 Water Report press release, 9 June 2026In the report, aging infrastructure tops a figure whose surrounding prose carries no percentage. The 66 percent belongs to a different question, on treatment as the top operational challenge created by new regulatory requirementsMore than 600 US water industry stakeholders, 2026Published, but the quoted form is a scope error
“Seventy-one percent of utilities cite staffing as a barrier”Same press release. Scoped form in the report: 71 percent of respondents whose digital-strategy objectives are not being metA follow-up question asked of a sub-populationA sub-population of that survey, 2026Published and traceable in scoped form only
“Human error causes 40% of data center outages”, “85% are procedure failures”Uptime Institute Annual Outage Analysis 2025, announced 6 May 2025. The 2026 edition, announced 13 May 2026, publishes no human-error percentageShare of organizations that had such an outage in three years, then a share of those human-error incidents. Neither is a share of outagesUptime Institute survey population, 2025Published for 2025, with inverted denominators in common use. Cannot be traced as a 2026 figure
“340,000 unfilled data center positions, per BLS”Industry reporting and recruiting commentary. The Bureau of Labor Statistics does not publish this figureNot establishedNot establishedCannot be traced to a primary
“Transformer lead times are 22 to 33 months”US Department of Energy, Office of Electricity, 22 February 2024, on distribution transformers only: “These same orders that previously took two to four months to fulfill now take 22 to 33 months”Average order-to-delivery time for distribution transformers across all voltage classesDOE commentary, 2024Published and traceable, but class-specific and stale. CRS Report R48933, 23 April 2026, records an industry-consultancy estimate of 30 weeks by Q2 2025
“Transformer lead time” quoted as one number for all equipmentWood Mackenzie Q2 2025 survey, in commentary of 15 October 2025: power transformers 128 weeks, generator step-up units 143 weeks, switchgear 44 weeksAverage quoted lead time by equipment classWood Mackenzie survey, Q2 2025Published and traceable, per class only
“95% of AI pilots fail”The GenAI Divide: State of AI in Business 2025, MIT NANDA, July 2025. Not MIT Sloan, not BCG95% of organizations getting zero return. Funnel for embedded or task-specific GenAI: 60% investigated, 20% piloted, 5% successfully implementedOver 300 disclosed initiatives, 52 interviews, 153 survey responses, January to June 2025Published and traceable, with the authors’ own limitation attached
“Emergency repairs drive most utility overtime”AWWA 2026 State of the Water Industry, Table 18, n=769How often each reason leads to overtime, on a never-to-always scale. On-call response 29.5% often plus 30.4% always. Emergency repairs 32.1% plus 22.1%Fielded 21 September to 31 October 2025, published 30 April 2026, total n=2,171. The 769 is a screened subgroup of utility respondents, not a response ratePublished and traceable, but a frequency of citation, not a share of overtime hours or cost. The table cannot rank drivers by hours
“About 10,700 annual openings for water and wastewater operators”The BLS 2024 to 2034 projection cycle, now replaced. BLS refreshed the occupation on 27 August 2026 to the 2025 to 2035 cycle: 131,000 falling to 123,500, minus 6%, about 10,000 openings a year, median wage $60,020 as of May 2025Occupational employment projectionBLS, 2025 base yearPublished and traceable, but the older figure is stale

Why do wrench-time benchmarks disagree?

Because the definitions disagree, and almost nobody publishes theirs.

The 2024 paper is an open primary measurement, and it shows how much structure sits under one percentage. Rahman sampled 320 maintenance activities over three months at one chemical processing plant and found an average of 28 percent across the site. Underneath that, the paper reports crew averages ranging from 20 to 35 percent and craft-specific wrench times ranging from 13.3 percent for one trade to 45.5 percent for another. Same site, same method, same three months.

The non-wrench time is the more useful half. The paper’s loss categories put permit and clearance at 11 percent, execution planning at 8 percent, waiting on process at 8 percent, looking for tools and PPE at 7 percent, looking for parts and material at 4 percent, and wrap-up activity at about 4 percent. Permits, planning, parts search and wrap-up alone come to roughly 27 percent of the observations. That is the shift, described.

Now put that next to a 55 percent target. If one organization counts travel to the job as work and another does not, if one books permit waiting inside the job and another to a separate code, the two numbers are not comparable. A single-plant study is not a population estimate for industry and does not claim to be. What it demonstrates is that the metric is definitional before it is empirical.

One more thing about that paper, and it is the reason it sits in this ledger rather than being quoted flat. It was published in 2024 and the observations were taken in the first half of 2023. Anyone citing “a 2024 wrench-time study” has already lost one of the four parts.

The two practitioners most associated with the metric warn against it themselves. Doc Palmer, where the 55 percent figure is usually traced to, wrote in Plant Services on 5 January 2022 under the headline that wrench time can be a terrible metric: “The craftsperson works directly at the pump all shift without ever checking in, needing extra parts, or even taking breaks. Additionally, the craftsperson has mistakenly taken the wrong pump out of service to repair. The person has 100% wrench time that day doing the wrong job in the wrong way!” Torbjörn Idhammar of IDCON has published six reasons to stop running wrench-time studies at all, the fifth of which is that “the product of a maintenance department is equipment reliability”, so a metric that scores thinking, planning and problem-solving as zero is measuring the wrong thing.

Treat wrench time as contested. Use it as a diagnostic for finding where the shift goes, never as a target, and read 55 percent quoted without a definition as telling you nothing.

Does the 10x predictive maintenance return exist?

Yes. This is the row people get wrong in the other direction.

The figures are in section 5.4, on page 5.4, of the Operations and Maintenance Best Practices Guide, Release 3.0, published August 2010 by Pacific Northwest National Laboratory for the Department of Energy’s Federal Energy Management Program. The guide is public and free. The passage reads: “In fact, independent surveys indicate the following industrial average savings resultant from initiation of a functional predictive maintenance program”, then lists return on investment 10 times, reduction in maintenance costs 25 to 30 percent, elimination of breakdowns 70 to 75 percent, reduction in downtime 35 to 45 percent, and increase in production 20 to 25 percent.

So anyone who says the numbers cannot be located is wrong, and will be corrected in any room with a federal energy manager in it. The defensible criticism is narrower and harder to answer. The guide does not identify the independent surveys. Chapter 5’s reference list contains two items, a NASA reliability-centred maintenance guide and a Pump-Zone article on proactive maintenance for pumps, and neither is a survey of predictive-maintenance savings. The trail ends at a 2010 secondary summary of unnamed work: the number and the year are readable, the population and the question are not.

Then the figures were recycled. They appear today across vendor blogs, analyst decks and procurement business cases as “the US Department of Energy found”, which is technically defensible and practically misleading, because 2010 predates the sensor economics, connectivity and software any current business case depends on. Cite them as what they are, and do not build a capital case on them alone.

What happened between the Black and Veatch report and its press release?

A clean worked example of the mechanism that produces most bad numbers in this market.

The Black and Veatch 2026 Water Report surveyed more than 600 US water industry stakeholders. The press release of 9 June 2026 says plainly: “Sixty-six percent of respondents cite aging infrastructure as a top challenge.” That sentence is now everywhere, usually compressed further into a claim about what most water utilities say.

In the report itself, aging infrastructure does top a figure, and the prose around it attaches no percentage. The 66 percent belongs to a different question, about treatment as the top operational challenge created by new regulatory requirements. The press release invented nothing. It broadened a scoped result into a general one, and the general one is more quotable.

The staffing figure shows the same broadening. The press release says “Seventy-one percent of utilities cite staffing as a barrier”. The report’s form is narrower: among respondents whose digital-strategy objectives are not being met, 71 percent cite staffing as a barrier and 48 percent cite legacy data and systems. The sub-population is the interesting part, because staffing is being named by people who have already said their digital programme is stalling. Strip the scope and you get a claim about utilities in general, which the report does not support.

The report also contains a figure needing no repair, and the press release states it correctly: “Seven in 10 (70%) say they collect sufficient data, but only 19% say they leverage it effectively.”

Which human-error percentage belongs to which year?

Both circulating percentages come from the Uptime Institute Annual Outage Analysis 2025, announced 6 May 2025. Neither is in the 2026 edition, announced 13 May 2026.

Both are also restated with the wrong denominator. The 2025 announcement says “Nearly 40% of organizations have suffered a major outage caused by human error over the past three years.” That is a share of organizations across a three-year window, not the share of outages attributable to human error, which is the form almost everyone quotes. The 85 percent is scoped inside that: “Of these incidents, 85% stem from staff failing to follow procedures or from flaws in the processes and procedures themselves.” It is not 85 percent of outages.

What the 2026 announcement contributes is qualitative and, for maintenance work, more useful than either percentage: “For 2026, failures to follow established procedures remain the leading driver of human error-related outages.” Attribute the percentages to 2025 and the ranking to 2026.

Why “transformer lead time” is never one number

Two different errors travel under this heading, and they need separating.

The first is a class error. Wood Mackenzie’s Q2 2025 survey, discussed in commentary published 15 October 2025 by Benjamin Boucher, Devin Thomas and Michael Mendrek-Laske, gives separate averages by equipment class: power transformers 128 weeks, down ten weeks on the prior quarter, generator step-up units 143 weeks, and switchgear 44 weeks. Those are not interchangeable. A repair plan that quotes any one of them as the category has averaged unrelated products.

The second is a currency error, and it is the one behind the 22-to-33-month figure still in circulation. That range is not a Wood Mackenzie number and does not cover four products. It comes from the Department of Energy’s Office of Electricity on 22 February 2024, it describes distribution transformers across all voltage classes, and the original sentence is comparative: “These same orders that previously took two to four months to fulfill now take 22 to 33 months.” That peak did not hold. Congressional Research Service Report R48933, published 23 April 2026, records that “an industry consultancy estimated that wait times for distribution transformers had decreased to 30 weeks by the second quarter of 2025”. CRS does not name the consultancy, which is worth saying out loud: the corrected figure is better than the stale one and is still a secondary attribution to an unnamed private source.

Note a small thing that matters on a page like this one. The generator step-up average appears as 143 weeks in the Wood Mackenzie commentary and comes back as 144 in secondary write-ups. Trivial in itself, and exactly the drift that becomes a range nobody can source. Name the class, name the date, or drop the number.

What the 95% AI pilot failure figure actually counted

The report is The GenAI Divide: State of AI in Business 2025, by Aditya Challapally, Chris Pease, Ramesh Raskar and Pradyumna Chari, published by MIT NANDA in July 2025. It is not an MIT Sloan publication and not a BCG study, and is routinely attributed to both.

The methodology is stated in the report: “a systematic review of over 300 publicly disclosed AI initiatives, structured interviews with representatives from 52 organizations, and survey responses from 153 senior leaders collected across four major industry conferences”, over a research period of January to June 2025. The authors attach their own limitation to the adoption funnel: “These figures are directionally accurate based on individual interviews rather than official company reporting. Sample sizes vary by category, and success definitions may differ across organizations.”

The funnel is more informative than the headline. The report charts two series. For embedded or task-specific GenAI, 60 percent of organizations investigated, 20 percent piloted, and 5 percent successfully implemented. For general-purpose LLMs the same steps read 80, 50 and 40 percent. The headline 95 percent is the complement of that first series’ last step and applies to that class of system, not to AI generally.

If you are gating a maintenance-AI purchase, the number to take from this report is not 95. It is that the drop happens between pilot and production. The report attributes it to a learning gap, and records users describing custom or vendor-pitched tools as “brittle, overengineered, or misaligned with actual workflows”.

FIG. 1

How the audited figures grade out

Scroll sideways to see the whole drawing.

Figure 1. How the audited figures grade out. A single horizontal bar divided into three proportional segments representing the verdicts reached in the ledger above. The first and smallest segment is figures that trace to a published primary with a stated method, sample and year. The second is practitioner benchmarks, which are widely published professional judgement rather than measured findings. The third and largest is figures that cannot be traced to any primary at all, or that trace to a source narrower than the claim made from it. The bar is proportional and unitless and counts rows in this ledger only.

Counted across the rows in the ledger above. The middle band is not a criticism: a practitioner benchmark is useful as long as nobody presents it as a measurement.

The practitioner benchmarks, and why that verdict is not a criticism

Ninety percent preventive maintenance compliance and one planner per 20 technicians are useful. They are the accumulated judgement of people who have run maintenance organizations, and a plant at 55 percent compliance with one planner for 60 technicians has a real problem these conventions will help it see.

The failure mode is citing them as findings. There is no population, no year and no question behind either. Palmer’s stated form is “one planner for every 20-30 craftspersons”, argued from planning economics rather than sampled from plants.

One conflation changes the number materially. PM compliance measures completion of preventive tasks inside a compliance window. Schedule compliance measures completion of all scheduled work, preventive and corrective, against the schedule that was committed. Different measurements, different denominators, routinely reported under each other’s names. Set a target against one and measure it against the other and it will be met or missed for reasons nobody in the room can explain.

EQUA’s own four numbers, graded by the same rule

We ran the ledger’s rule on ourselves. Here is what it cost us.

EQUA AI publishes four figures. All four come from one anonymized industrial production deployment, live since January 2026, with 10 critical assets, 5 active production users, a daily operating cadence and approval-gated workflows. That is the entire population. Not a study, not a multi-site result, not a water-sector result, and not a forecast for anyone else’s site.

EQUA figureWhat was measuredSample and periodVerdict by the same rule
31.7% lower MTTR, 14.2 hours to 9.7 hoursMean time to repair on the deployed repair workflow. Time to repair, not total downtime, which is a larger and different quantityOne deployment, 10 critical assets, 5 users, since January 2026Internal measurement, method not published in full
80.4% faster quote cycles, 4.6 days to 0.9 daysElapsed time on the deployed quoting workflowSameInternal measurement, method not published in full
96% faster quote-to-order, 3 days to less than 1 hourThe same interval on the deployed quote-to-order workflow, on a business-hours clock rather than a calendar clockSameInternal measurement, clock-dependent, method not published in full
50% lower measured safety-risk exposureA risk-exposure measure on workflows with defined approval gatesSameInternal measurement, method not published in full

The clock is the point. Which clock you run changes the percentage for an identical stretch of wall time, because a business-hours clock does not count nights and weekends and a calendar clock does. Quoting the 96 percent without naming the clock would be the same scope error this article objects to elsewhere. So the clock travels with the figure, every time, including here where it makes the number look smaller.

The measurement method behind these four figures has not been published in full. On the four-part test: the number, the year and the population are readable, and the population is small. The exact question is only partly readable, because the workflow boundaries and exclusions are described but not documented to the standard this page demands of everyone else. Weight them accordingly. Ask for the protocol before treating any of them as a forecast for your own site, and expect a straight answer. EQUA has no named customer and no published case study, and no number above is presented as one.

Who wrote this

EQUA AI builds EQUA AIMMS. On every fault it touches it keeps a record of what it read, when it read it, what it concluded and who it told.

That is the difference between a number you can put in front of a board and a number you have to defend from memory. Most maintenance figures fail the audit above because nothing was recording the denominator while the work happened. If you are going to hold vendors to this standard, hold your own operating data to it, and run something that can produce the answer.

Sources

Turn this idea into a facility-specific decision.

Bring one recurring failure or stuck workflow. The path is deliberately focused:

  1. 01

    Intake

    Complete a short qualification intake.

  2. 02

    Working session

    Map the delay and control boundary in 20 minutes.

  3. 03

    First-scope decision

    Decide whether a credible facility-specific first scope exists.

Book Your 20-Minute Assessment