DataCentreNews Ireland - Specialist news for cloud & data centre decision-makers
Ireland
Five control rooms, one tuesday afternoon

Five control rooms, one tuesday afternoon

Mon, 27th Jul 2026 (Today)
Kamalensh Srikanth
KAMALENSH SRIKANTH Product Strategy Leader AlertOps

Picture five control rooms on the same hot afternoon.

In a wholesale colocation data center, a chilled-water plant trips on a primary loop. Within ninety seconds the building management system has issued forty-seven alarms across two halls, and air temperature is already climbing toward the racks of tenants who pay for Tier III uptime.

In a carrier network operations center, a backhoe two states away severs a fiber. Tens of thousands of circuits go dark at once, and the screen fills with link-down alarms from every customer riding that path.

On an automotive assembly line, a robot cell faults. The PLC throws a burst of tag alarms, the Andon signal fires, and the line stops at a cost measured in dollars per second.

On a midstream oil and gas pipeline, pressure drops on an unmanned segment. SCADA flags an anomaly that could be a sensor glitch or a leak with environmental and safety consequences.

At an electric utility, a substation breaker trips and remote terminal units light up across the territory. Any one of them could be the cause or just a symptom downstream.

Five industries. Five completely different machines. Underneath, the same event: the physical world just failed, and a wall of alarms is demanding a human response measured in seconds, not business days.

This is mission-critical operations. The discipline for handling these moments has a name: Mission-Critical Incident Management (MCIM). The equipment differs wildly, but the anatomy of the incident, from first alarm to closed audit record, is remarkably consistent. Learn the anatomy once, and you can read any of these control rooms.

Why a Mission-Critical Incident Is a Different Animal

A mission-critical incident is not a louder version of an IT incident. It is a different category entirely, and the distinction matters before you can manage one well.

The consequences are physical. An IT incident, at worst, loses or exposes data. A mission-critical incident can destroy equipment worth millions, release hazardous material into the environment, cut service to a region, or put someone in physical danger. Severity is not a queue label. It is a statement about human and equipment risk.

The priorities are inverted. IT disciplines teach confidentiality first. Operational technology flips that order: safety first, then availability, then integrity, with confidentiality last. An operator would rather someone read their process data than take a safety system offline.

The constraints are unforgiving. You do not reboot a running turbine or patch a live refinery. Maintenance happens in planned windows, sometimes a year apart. Equipment runs for decades. Physics does not wait for a triage queue. A cooling loss or pressure excursion can cascade in seconds.

The responders are not at a desk. They are shift operators in control rooms, field technicians on catwalks or in switchyards, on-call engineers, and plant managers. Reaching the right person is half the battle.

Regulators are in the room. NERC CIP, process-safety standards, environmental reporting requirements: these shape what resolved even means. An incident is not closed until an auditable record exists.

That last point defines why the anatomy runs from first alarm all the way to the audit record. Restoring service is necessary, but it is not sufficient.

The Six Stages: A Universal Anatomy 

Stage 1: Detection

Every incident starts with a trigger, typically an alarm from a monitoring or control system. The trap is that one physical failure almost never produces one alarm.

The chiller trip generates forty-seven BMS alarms as every downstream pump, air handler, and temperature sensor reports the loss in sequence. The fiber cut generates thousands of loss-of-signal events. The robot fault creates a burst of PLC tags. The substation event ripples across RTUs.

This is not a hypothetical nuisance. It is a measured, governed problem. The alarm-management standards ANSI/ISA-18.2 and the EEMUA 191 guideline define an acceptable rate for a single operator as fewer than one alarm per ten minutes, with a flood defined as more than ten in ten minutes. Forty-seven alarms in ninety seconds is not an inconvenience against those benchmarks. It is a catastrophic failure of signal-to-noise.

The detection challenge in mission-critical operations is rarely whether we saw it. It is whether a human can find the one signal that matters inside the storm.

Stage 2: Correlation

This is the hinge of the entire anatomy. Correlation groups related alarms by physical cause, shared path, and asset hierarchy into a single incident with a single owner, and enriches it with context: which asset, which location, which tenant or customer, what severity.

The effect is consistent across every industry. Forty-seven BMS alarms become one cooling incident mapped to Hall 2. Ten thousand loss-of-signal alarms become one fiber cut on a known route. A hundred process alarms become one unit upset. Once the flood collapses to a root cause, the operator works the cause instead of chasing fifty symptoms, and the metric that matters most, mean time to acknowledge (MTTA), drops from minutes to seconds.

Correlation is also what makes everything downstream possible. You cannot route to the right person, notify the right stakeholder, or write a clean record until you know what actually happened.

Stage 3: Escalation and Routing

A correlated incident is worthless if it pages the wrong person, reaches the right person at the wrong time, or goes unacknowledged because that person is forty feet up a cooling tower.

Mission-critical routing must be multi-dimensional: by skill, shift, and location, across every channel a responder might be reachable on, including mobile push, SMS, email, chat, and voice fallback when a page goes unacknowledged. The right responder is industry-specific, but the logic is universal. A chilled-water fault needs the on-call mechanical engineer, not the nearest available person. A fiber cut needs a transport technician. A pipeline alarm needs a field operator who can physically reach an unmanned site.

Routing by skill and location rather than a flat on-call list is often the difference between a two-minute and a forty-minute response. Because cascades move in seconds, escalation must be automatic and time-bound. If no one acknowledges, the incident escalates and the phone rings. There is no queue for a thermal runaway.

Stage 4: Stakeholder Notification

While responders work the cause, a parallel obligation fires: the people affected by the incident need to be informed accurately, quickly, and in the right words.

The audiences vary by industry; the pattern does not. A colocation operator owes proactive notice to Tier III tenants and a different message to lower tiers. A carrier owes affected customers real scope and a credible ETA, not ten conflicting updates. A plant owes production, planning, and quality a heads-up before bad product ships. An energy operator owes HSE, partners, and regulators notice the moment a reportable threshold is crossed.

There are usually two possible messages: handled, you are safe versus at risk, here is what we are doing. Sending the wrong one destroys credibility that takes years to rebuild. Pre-positioned templates, fired automatically by severity and stakeholder tier, transform a frantic manual scramble into coordinated, defensible communication. This also pre-empts the flood of inbound calls from customers and tenants asking whether their service is okay, calls that would otherwise consume the very responders trying to fix the problem.

Stage 5: Resolution Against Procedure

In mission-critical environments, you do not improvise on live plant. Resolution is governed by an approved Method of Procedure (MOP), Standard Operating Procedure (SOP), or Emergency Operating Procedure (EOP): the chiller-failure MOP, the grid-loss EOP, the line-restart SOP.

Effective incident handling links that procedure to the live incident, enables the team to collaborate in their own channels, and logs each step with a timestamp as it is completed. The IT service management system of record, typically ServiceNow or Jira, stays authoritative and synced bidirectionally so nothing lives in two disconnected places. The procedure is followed and evidenced at the same time, which sets up the final and most frequently overlooked stage.

Stage 6: The Audit Record

In every one of these industries, service restored is only half the job. The incident is not truly closed until the record exists.

A colocation operator needs a per-tenant timeline for a Tier audit or SLA-credit dispute. A utility needs evidence for a NERC CIP review. An energy operator needs a process-safety-grade account of what happened and who authorized what. A manufacturer needs accurate downtime reason codes to protect OEE. Reconstructing any of this by hand, after the fact and under pressure, is slow, error-prone, and exactly what auditors distrust.

The discipline is to make the audit-ready record a by-product of the response itself, not a four-hour homework assignment afterward. Every alarm, escalation, notification, and procedure step is captured chronologically in real time. The post-incident review then feeds back into the system, tuning correlation rules so the next flood is smaller and refining routing so the next page lands faster. The anatomy closes the loop.

The Boundary That Earns Credibility

One principle runs through all six stages and is worth stating plainly, because it is what earns trust with operators: orchestration happens above the control layer. The system that correlates, routes, notifies, and documents does not actuate equipment and is not a safety system. The SCADA, BMS, and DCS retain control. The safety instrumented system keeps the process safe. Incident management coordinates the human response on top of all of it.

Operators will not trust any approach that blurs that line. They will immediately trust one that states it clearly, and states it first.

Why the Anatomy Matters

Walk back through the five control rooms. The chiller, the fiber, the robot, the pipeline, the substation could not be more different. But the response that separates a rehearsed team from a chaotic one is identical in shape: detect, correlate, route, notify, resolve, document, above the control layer, never touching the equipment.

The machines are industry-specific. The motion is universal.

Organisations that treat this anatomy as a discipline, and instrument it rather than improvise it, transform their worst day into something closer to a drill: an alarm flood that collapses to one incident, reaches the right person in under two minutes, keeps every stakeholder correctly informed, follows the approved procedure, and writes its own audit trail on the way out.

The first alarm is going to fire no matter what. What you do in the minutes that follow, and the record you can produce when it is over, is the part you actually control.