A chiller drops and N+1 becomes N. The clock that matters starts there.
AIMMS handles the permitted digital tasks around high-consequence facility repairs. It connects alarms, procedures, asset history, redundancy context, exact parts, cross-site spares, suppliers, warranties, approvals, maintenance windows, and closeout.
What you get Redundancy state, open work, the correct MOP, the exact spare and the vendor position in one place while you still have margin to spend.
- No BMS or EPMS writes
- Redundancy state stays operator-owned
- Change windows respected
Read-only facility context · Customer-defined approvals · Direct facility-control writes permanently out of scope
The industry’s own outage research keeps landing on procedure and process, not equipment.
Each one is a published third-party figure from the Uptime Institute outage research programme or the North American Electric Reliability Corporation, carrying its report and its year. None is an EQUA result and none describes your estate. What each one does to a repair, and what AIMMS does about that, is stated beside it.
-
The failure is almost never in the compute.
45% of impactful outages are power failures, with cooling accounting for a further 14% Uptime Institute, Annual Outage Analysis 2026, opens in a new tabDuring the repairAvailability depends on switchgear, UPS, batteries, generators, chillers, cooling towers and pumps, which are maintained assets with parts, suppliers, procedures and approvals behind every repair.
What AIMMS changesAssembles the evidence around a fault in the supporting mechanical and electrical plant with each item traced to its source, so the exposure window is governed by the work rather than by the coordination around it.
-
Almost all of it involves a person.
92% of operators say human error was at least a minor contributor to their significant outages of the past three years Uptime Institute, Annual Outage Analysis 2026, opens in a new tabDuring the repairThe people carrying that risk are the same people reconstructing the context around it: pulling the redundancy state, finding the change record, and confirming which unit is actually isolated, under time pressure.
What AIMMS changesCarries the case state, the evidence it used, the decision and the approval state on every action, so the next person picks up a record rather than a recollection.
-
And almost all of that is the procedure.
85% of human-error outages come from staff not following procedures, or from flaws in the procedures themselves Uptime Institute, Annual Outage Analysis 2025, opens in a new tabDuring the repairThis is not a competence finding. It locates the failure in whether the right procedure, in its right revision, reached the right person at the right moment, which is a retrieval and coordination problem.
What AIMMS changesKeeps procedures, as-built evidence, asset configuration and verified incident lessons attached to the equipment instead of ageing in a handover folder nobody opens.
-
There is less room for downtime.
224 GW of additional summer peak demand forecast across North America over ten years, roughly 69% above the previous year’s forecast, with new data-center load accounting for most of the increase NERC 2025 Long-Term Reliability Assessment, opens in a new tabDuring the repairCapacity is being committed faster than the plant supporting it can be maintained, so a redundancy-constraining event that used to be absorbed now competes with a schedule that has no slack in it.
What AIMMS changesSurfaces which dependencies are still unsatisfied around a maintenance window and prepares the coordination tasks against them, so readiness is visible before the window rather than discovered inside it.
Together these locate the exposure in the supporting physical plant and in whether the right information reached the right person while the clock was running. Neither is an equipment problem, and neither is solved by more monitoring.
Published third-party sector research does not establish AIMMS effectiveness. Only a baselined deployment on the customer's own workflow can do that.
Price the hours you spend at reduced redundancy waiting on a decision.
Two numbers from your own incident record: how many cooling or critical-power events constrain redundancy in a year, and how long each one holds between the alarm and an approved, executable work path. No projection is applied.
Move both controls. Use your own incident record, not our example.
These are your incident-record numbers. No restored margin is credited to AIMMS anywhere on this page. Nothing you type here leaves your browser.
This section needs JavaScript to do your arithmetic. The figures beside it are the worked example for 16 redundancy-constraining events a year at 11 waiting hours each. To use your own numbers, open the full value model.
176 hours a year
That is time at reduced margin spent assembling a decision, not executing one.
SLA exposure, deferred capacity, vendor premium, escalation load, insurance position. Enter the number your risk register already uses, or leave it empty.
Add an hourly figure and this becomes an annual exposure your risk register and your finance team can both check.
They arrive there as events a year, avoidable hours per event, and cost per hour, where your finance team can argue with each one separately.
Book Your 20-Minute AssessmentProtect uptime, and prove it on your own plant.
Four operating outcomes the AIMMS workflow is built to produce around a cooling or critical-power event. They are outcomes we design for and draw hatched.
Your KPI contract On a facility pilot, the deciding numbers come off your own incident record and window log. Six KPIs, agreed and baselined on the selected cooling or critical-power workflow before the first incident is routed through AIMMS, then reported against that baseline. See the KPI contract this pilot will measure.
Build the facility value model on your own risk and SLA economicsMeasure the time an incident spends waiting on coordination.
Baseline the current incident and window workflow before the pilot starts, measure the pilot against that baseline, and expand only when the evidence supports the decision. Every number here comes off your own plant.
- Mean time to ready work
- Elapsed time from accepted alarm to an approved, executable path.
- Incident duration
- Time the selected asset or system remains constrained or unavailable.
- Redundancy-context readiness
- Required current constraints and open work available before the decision.
- Part certainty
- Exact variant, approved alternative, and usable local or cross-site stock confirmed.
- Vendor and warranty cycle
- Elapsed time to a usable external response or authorization.
- Window and closeout quality
- Required work, approvals, evidence, and final record complete.
A data center incident crosses systems faster than its operating context does.
A single event can touch cooling capacity, electrical redundancy, generator readiness, water treatment, warranties, maintenance windows, spares, contractors, and customer commitments. Each team often starts with only one fragment.
Where elapsed time accumulates
- 01 Alarm and impact identified
- 02 Redundancy and open work checked
- 03 Procedure and evidence assembled
- 04 Exact part and cross-site stock verified
- 05 Vendor, warranty, and change approval moved
- 06 Work completed and operating memory updated
AIMMS creates one governed action path across power, cooling, water, facilities, suppliers, and approvals.
Modelled clock from the alarm to restored redundancy
11 of 17 modelled hours are spent on a thinner margin than the design intent, with nobody at the equipment.
- 0.5 h Alarm raised, redundancy state checked Waiting on coordination
- 2.5 h Procedure, open work and vendor position assembled Waiting on coordination
- 3 h Exact variant and cross-site stock verified Waiting on coordination
- 3 h Vendor commitment and change approval Waiting on coordination
- 5 h Work inside the change window Hands on the asset
- 1 h Verify and restore redundancy Verified and returned
- 2 h Closeout and commissioning record updated Waiting on coordination
Segment lengths follow the modelled premise stated further down this page. Hatched blocks are modelled waiting, the solid block is the work inside the window, and the green block is the operator’s own restored redundancy.
What reduced-redundancy waiting is worth across one campus in a year.
A 12 MW campus taking 16 cooling or critical-power events a year that constrain redundancy, each holding 11 hours between the alarm and an approved, executable work path.
- 16 redundancy-constraining events a year
- Cooling and critical-power support only. Not every ticket, and not IT or network incidents.
- 11 h waiting per event
- From alarm raised to a path with redundancy state confirmed, the procedure identified, the spare located, the vendor committed and the change approved. Excludes the work in the window and the restoration.
- 0 h improvement assumed
- No AIMMS effect appears anywhere in this arithmetic. The model sizes the exposure and stops.
176 modelled hours a year at reduced redundancy
Roughly 7 elapsed days a year on a thinner margin than the design intent, spent assembling a decision.
16 events multiplied by 11 waiting hours is 176 hours. Each input is an assumption we chose and printed, which is why the figure is hatched rather than solid. Replace all three with your own in the calculator at the top of this page.
Modelled from the stated premise and the assumptions shown. No AIMMS effect is applied.
Start with facility systems where coordination delay can consume operating margin.
Choose one critical asset family and one repeatable incident or maintenance-window workflow with clear redundancy, change, and authority rules.
Cooling
Chillers, pumps, cooling towers, heat exchangers, VFDs, CRAH or CRAC support, and thermal-management auxiliaries.
Critical power support
Generators, ATS, UPS and battery support, switchgear auxiliaries, station power, fuel systems, and electrical cooling support.
Water and treatment
Cooling-water treatment, reuse systems, pumps, chemical feed, blowdown, wastewater conveyance, membranes, and instrumentation.
Operating readiness
MOPs, SOPs, EOPs, commissioning evidence, deficiencies, critical spares, warranties, vendor cases, and maintenance windows.
What does not change when the asset does.
The parts above are specific to data centers & critical facilities. These are not: they are the same system, the same boundary and the same business case whichever operation you run.
- The operating contract The same seven questions answered at all six stages: what starts it, what AIMMS reads, what it does, who decides, what happens when something is missing, what is left behind, and what a pilot measures.
- One fault, followed end to end A single pump fault from alarm to verified return, and the Data Twin that gives AIMMS approved context without replacing a system of record.
- Security, control and authority Who can do what, what is read-only, what AIMMS does when a dependency fails, and why no setting exists that would let it write to a control system.
- The business case, on your numbers Five visible formulas, every input supplied by your finance team, and our measured production evidence kept on the other side of an explicit boundary.
Direct answers for critical facilities, portfolio operations, risk, and IT.
The first conversation should resolve which incident workflow to pilot, how far the boundary sits from BMS and EPMS, and how redundancy exposure will be baselined.
Does AIMMS replace our DCIM, BMS, EPMS, CMMS, commissioning, or purchasing systems?
No mandatory replacement is required. Existing systems remain authoritative. AIMMS creates a controlled Data Twin that joins permitted context around the selected fault-to-fix workflow and coordinates the digital work between systems and teams. Pilot writeback stays inside AIMMS. External CMMS, EAM, or ERP writeback requires customer approval and a validated integration.
Can AIMMS change BMS or EPMS setpoints, switch equipment, or start generators?
No. BMS and EPMS setpoint changes, switching, generator commands, redundancy-state changes, fire-system actions, and other live-control writes are permanently outside AIMMS scope. Facility personnel retain operating, change, safety, and return authority.
Which data-center workflows are strongest for a first pilot?
Strong starting points include chiller or cooling-pump response, generator or ATS support, cooling-water treatment, critical-spares assurance, commissioning-deficiency closure, and maintenance-window readiness.
How does the Data Twin help after commissioning?
It keeps the selected as-built configuration, test evidence, procedures, deficiencies, asset history, spares, suppliers, warranties, approvals, and verified incident outcomes connected as a living operating context.
What is the target timeline for a first facility pilot?
The working target is two weeks to configure the Data Twin and workflow, followed by a focused 30-day pilot across 10 to 20 selected assets. Scope and timing adjust to data readiness, security review, change rules, and integration complexity.
Where do the Uptime Institute figures on this page come from?
The benchmark band cites the Uptime Institute reports inline. The stated-premise model prints every assumption beside the hatched figure. The measured deployment scope and limitations sit directly below the four EQUA metrics.
Bring us the cooling, power-support, water, or handoff workflow that keeps creating risk.
In 20 minutes, we will map the incident path, the permitted digital work AIMMS can move, the facility authority boundary, and a measurable first-deployment target.
One operating problem. One focused working session. No obligation.