Secondary sector · Data Centers & Critical Facilities

A chiller drops and N+1 becomes N. The clock that matters starts there.

AIMMS handles the permitted digital tasks around high-consequence facility repairs. It connects alarms, procedures, asset history, redundancy context, exact parts, cross-site spares, suppliers, warranties, approvals, maintenance windows, and closeout.

What you get Redundancy state, open work, the correct MOP, the exact spare and the vendor position in one place while you still have margin to spend.

  • No BMS or EPMS writes
  • Redundancy state stays operator-owned
  • Change windows respected
data-center mechanical room with cooling pumps, piping, and electrical equipment
Data centers and critical facilities

Read-only facility context · Customer-defined approvals · Direct facility-control writes permanently out of scope

Why maintenance is getting harder here

The industry’s own outage research keeps landing on procedure and process, not equipment.

Each one is a published third-party figure from the Uptime Institute outage research programme or the North American Electric Reliability Corporation, carrying its report and its year. None is an EQUA result and none describes your estate. What each one does to a repair, and what AIMMS does about that, is stated beside it.

  1. The failure is almost never in the compute.

    45% of impactful outages are power failures, with cooling accounting for a further 14% Uptime Institute, Annual Outage Analysis 2026, opens in a new tab
    During the repair

    Availability depends on switchgear, UPS, batteries, generators, chillers, cooling towers and pumps, which are maintained assets with parts, suppliers, procedures and approvals behind every repair.

    What AIMMS changes

    Assembles the evidence around a fault in the supporting mechanical and electrical plant with each item traced to its source, so the exposure window is governed by the work rather than by the coordination around it.

  2. Almost all of it involves a person.

    92% of operators say human error was at least a minor contributor to their significant outages of the past three years Uptime Institute, Annual Outage Analysis 2026, opens in a new tab
    During the repair

    The people carrying that risk are the same people reconstructing the context around it: pulling the redundancy state, finding the change record, and confirming which unit is actually isolated, under time pressure.

    What AIMMS changes

    Carries the case state, the evidence it used, the decision and the approval state on every action, so the next person picks up a record rather than a recollection.

  3. And almost all of that is the procedure.

    85% of human-error outages come from staff not following procedures, or from flaws in the procedures themselves Uptime Institute, Annual Outage Analysis 2025, opens in a new tab
    During the repair

    This is not a competence finding. It locates the failure in whether the right procedure, in its right revision, reached the right person at the right moment, which is a retrieval and coordination problem.

    What AIMMS changes

    Keeps procedures, as-built evidence, asset configuration and verified incident lessons attached to the equipment instead of ageing in a handover folder nobody opens.

  4. There is less room for downtime.

    224 GW of additional summer peak demand forecast across North America over ten years, roughly 69% above the previous year’s forecast, with new data-center load accounting for most of the increase NERC 2025 Long-Term Reliability Assessment, opens in a new tab
    During the repair

    Capacity is being committed faster than the plant supporting it can be maintained, so a redundancy-constraining event that used to be absorbed now competes with a schedule that has no slack in it.

    What AIMMS changes

    Surfaces which dependencies are still unsatisfied around a maintenance window and prepares the coordination tasks against them, so readiness is visible before the window rather than discovered inside it.

Together these locate the exposure in the supporting physical plant and in whether the right information reached the right person while the clock was running. Neither is an equipment problem, and neither is solved by more monitoring.

Published third-party sector research does not establish AIMMS effectiveness. Only a baselined deployment on the customer's own workflow can do that.

Your own arithmetic

Price the hours you spend at reduced redundancy waiting on a decision.

Two numbers from your own incident record: how many cooling or critical-power events constrain redundancy in a year, and how long each one holds between the alarm and an approved, executable work path. No projection is applied.

Move both controls. Use your own incident record, not our example.

These are your incident-record numbers. No restored margin is credited to AIMMS anywhere on this page. Nothing you type here leaves your browser.

This section needs JavaScript to do your arithmetic. The figures beside it are the worked example for 16 redundancy-constraining events a year at 11 waiting hours each. To use your own numbers, open the full value model.

Your annual coordination exposure 16 × 11 h

176 hours a year

That is time at reduced margin spent assembling a decision, not executing one.

SLA exposure, deferred capacity, vendor premium, escalation load, insurance position. Enter the number your risk register already uses, or leave it empty.

Add an hourly figure and this becomes an annual exposure your risk register and your finance team can both check.

Carry these numbers into the full value model

They arrive there as events a year, avoidable hours per event, and cost per hour, where your finance team can argue with each one separately.

Book Your 20-Minute Assessment
What the operator gets

Protect uptime, and prove it on your own plant.

Four operating outcomes the AIMMS workflow is built to produce around a cooling or critical-power event. They are outcomes we design for and draw hatched.

One operating picture, not five fragments Redundancy state, open work, the procedure, the spare and the vendor position sit together, so facilities, electrical and vendor teams stop working from different halves of the incident. Operating outcome
Maintenance windows get used, not spent getting ready Parts, evidence, vendor commitments and approvals are complete before the window opens, so the window does the work it was booked for. Operating outcome
Commissioning knowledge survives the handover As-built configuration, test evidence, deficiencies and warranties stay attached to the equipment and keep paying back at every later incident. Operating outcome
Redundancy and change authority stay with the operator No BMS or EPMS setpoint writes, no switching, no generator starts, no redundancy-state changes. Facility personnel retain change-window and return-to-service authority. Outside AIMMS scope

Your KPI contract On a facility pilot, the deciding numbers come off your own incident record and window log. Six KPIs, agreed and baselined on the selected cooling or critical-power workflow before the first incident is routed through AIMMS, then reported against that baseline. See the KPI contract this pilot will measure.

Build the facility value model on your own risk and SLA economics
KPI contract

Measure the time an incident spends waiting on coordination.

Baseline the current incident and window workflow before the pilot starts, measure the pilot against that baseline, and expand only when the evidence supports the decision. Every number here comes off your own plant.

Mean time to ready work
Elapsed time from accepted alarm to an approved, executable path.
Incident duration
Time the selected asset or system remains constrained or unavailable.
Redundancy-context readiness
Required current constraints and open work available before the decision.
Part certainty
Exact variant, approved alternative, and usable local or cross-site stock confirmed.
Vendor and warranty cycle
Elapsed time to a usable external response or authorization.
Window and closeout quality
Required work, approvals, evidence, and final record complete.
Cooling alarm · Redundancy constrained

A data center incident crosses systems faster than its operating context does.

A single event can touch cooling capacity, electrical redundancy, generator readiness, water treatment, warranties, maintenance windows, spares, contractors, and customer commitments. Each team often starts with only one fragment.

Where elapsed time accumulates

  1. 01 Alarm and impact identified
  2. 02 Redundancy and open work checked
  3. 03 Procedure and evidence assembled
  4. 04 Exact part and cross-site stock verified
  5. 05 Vendor, warranty, and change approval moved
  6. 06 Work completed and operating memory updated

AIMMS creates one governed action path across power, cooling, water, facilities, suppliers, and approvals.

Chiller loop fault with redundancy at risk The load is unchanged. The margin that was covering it is not.

Modelled clock from the alarm to restored redundancy

11 of 17 modelled hours are spent on a thinner margin than the design intent, with nobody at the equipment.

  1. 0.5 h Alarm raised, redundancy state checked Waiting on coordination
  2. 2.5 h Procedure, open work and vendor position assembled Waiting on coordination
  3. 3 h Exact variant and cross-site stock verified Waiting on coordination
  4. 3 h Vendor commitment and change approval Waiting on coordination
  5. 5 h Work inside the change window Hands on the asset
  6. 1 h Verify and restore redundancy Verified and returned
  7. 2 h Closeout and commissioning record updated Waiting on coordination

Segment lengths follow the modelled premise stated further down this page. Hatched blocks are modelled waiting, the solid block is the work inside the window, and the green block is the operator’s own restored redundancy.

Stated-premise model

What reduced-redundancy waiting is worth across one campus in a year.

A 12 MW campus taking 16 cooling or critical-power events a year that constrain redundancy, each holding 11 hours between the alarm and an approved, executable work path.

16 redundancy-constraining events a year
Cooling and critical-power support only. Not every ticket, and not IT or network incidents.
11 h waiting per event
From alarm raised to a path with redundancy state confirmed, the procedure identified, the spare located, the vendor committed and the change approved. Excludes the work in the window and the restoration.
0 h improvement assumed
No AIMMS effect appears anywhere in this arithmetic. The model sizes the exposure and stops.
One hatched tile is one modelled event. Hatched fill marks the stated-premise model.

176 modelled hours a year at reduced redundancy

Roughly 7 elapsed days a year on a thinner margin than the design intent, spent assembling a decision.

16 events multiplied by 11 waiting hours is 176 hours. Each input is an assumption we chose and printed, which is why the figure is hatched rather than solid. Replace all three with your own in the calculator at the top of this page.

Modelled from the stated premise and the assumptions shown. No AIMMS effect is applied.

Start where the margin is thinnest

Start with facility systems where coordination delay can consume operating margin.

Choose one critical asset family and one repeatable incident or maintenance-window workflow with clear redundancy, change, and authority rules.

data-center cooling plant with chilled-water piping, pumps, and headers
Cooling plant: where the redundancy margin is spent

Cooling

Chillers, pumps, cooling towers, heat exchangers, VFDs, CRAH or CRAC support, and thermal-management auxiliaries.

Critical power support

Generators, ATS, UPS and battery support, switchgear auxiliaries, station power, fuel systems, and electrical cooling support.

Water and treatment

Cooling-water treatment, reuse systems, pumps, chemical feed, blowdown, wastewater conveyance, membranes, and instrumentation.

Operating readiness

MOPs, SOPs, EOPs, commissioning evidence, deficiencies, critical spares, warranties, vendor cases, and maintenance windows.

Data Centers & Critical Facilities FAQ

Direct answers for critical facilities, portfolio operations, risk, and IT.

The first conversation should resolve which incident workflow to pilot, how far the boundary sits from BMS and EPMS, and how redundancy exposure will be baselined.

Does AIMMS replace our DCIM, BMS, EPMS, CMMS, commissioning, or purchasing systems?

No mandatory replacement is required. Existing systems remain authoritative. AIMMS creates a controlled Data Twin that joins permitted context around the selected fault-to-fix workflow and coordinates the digital work between systems and teams. Pilot writeback stays inside AIMMS. External CMMS, EAM, or ERP writeback requires customer approval and a validated integration.

Can AIMMS change BMS or EPMS setpoints, switch equipment, or start generators?

No. BMS and EPMS setpoint changes, switching, generator commands, redundancy-state changes, fire-system actions, and other live-control writes are permanently outside AIMMS scope. Facility personnel retain operating, change, safety, and return authority.

Which data-center workflows are strongest for a first pilot?

Strong starting points include chiller or cooling-pump response, generator or ATS support, cooling-water treatment, critical-spares assurance, commissioning-deficiency closure, and maintenance-window readiness.

How does the Data Twin help after commissioning?

It keeps the selected as-built configuration, test evidence, procedures, deficiencies, asset history, spares, suppliers, warranties, approvals, and verified incident outcomes connected as a living operating context.

What is the target timeline for a first facility pilot?

The working target is two weeks to configure the Data Twin and workflow, followed by a focused 30-day pilot across 10 to 20 selected assets. Scope and timing adjust to data readiness, security review, change rules, and integration complexity.

Where do the Uptime Institute figures on this page come from?

The benchmark band cites the Uptime Institute reports inline. The stated-premise model prints every assumption beside the hatched figure. The measured deployment scope and limitations sit directly below the four EQUA metrics.

Bring us the cooling, power-support, water, or handoff workflow that keeps creating risk.

In 20 minutes, we will map the incident path, the permitted digital work AIMMS can move, the facility authority boundary, and a measurable first-deployment target.

One operating problem. One focused working session. No obligation.