Reliability Friday 03: How to Build an Incident Escalation Matrix
Downtime averages $4,537 a minute. You pay that for every one of the fifteen minutes spent finding out who approves a rollback. Test your matrix with a seven-line template.

Short answer: Reliability does not come from tooling alone. If the person you hold responsible has no authority to decide, the incident drags on waiting for approval, not for a technical fix. And that wait has a measurable cost: according to PagerDuty's 2026 data, 68% of organizations lose more than $300,000 an hour during IT incidents, and 34% lose at least $500,000. The same research puts the average cost of downtime at $4,537 per minute and the average time to resolve an incident at 175 minutes.
Test your escalation matrix against the seven lines below: every line left blank is time you will lose during a real outage.
Don't Deploy on Friday. Verify on Friday.
This week's check
There is no command to run this week. Open your team's escalation matrix and try to fill in these seven lines:
Service: ?
First called if Severity 1 fires: ?
Who decides on rollback: ?
Who decides on failover: ?
Who informs the customer: ?
Does on-call have prod access (Y/N): ?
Date of last drill / post-mortem: ?A line whose answer is "it depends" is a line that has not been filled in. Asking that question for the first time during an outage means minutes spent looking for the answer. At $4,537 a minute, fifteen minutes spent tracking down rollback authority costs roughly $68,000, and no engineer writes a line of code in that time.
Without a severity definition the matrix does not work
The part of the matrix most often skipped is not who gets called but what counts as Severity 1. If the definition is not written down, the first five minutes go to debating "is this a Sev 1 or a Sev 2?" instead of investigating. Base it on simple, measurable criteria such as customer impact, the share of users affected, and whether revenue has stopped; an adjective like "important" is too vague to act on.
The checklist
Once the matrix is filled in, test it with eight questions:
- Who takes the first action during an incident?Who holds decision authority?Who can decide on a rollback?Who can decide on a failover?Who manages customer communication?Is the escalation order clear?Does the on-call team have the access it needs?Is there a post-mortem process?
The seventh is often overlooked. Holding on-call engineers responsible without giving them metrics, tools and access is a common mistake: the person is paged, sees the problem, and cannot act.
The eighth is one everyone agrees with and few actually do. In the same research, all surveyed organizations say post-incident learning is necessary, yet only 48% turn incidents into structured learning. That gap is where you find the teams that live through the same outage every quarter.
Expected output
A completed Severity 1 row looks roughly like this. It is an example; the roles depend on your organization:
Severity 1 incident
Owner: NOC / Operations
Technical Lead: Infrastructure Lead
Business Owner: Account / Service Owner
Decision Rights: rollback and failover authority defined
Communication: Customer Success + Service Desk
Post-mortem: mandatory within 48 hoursThe job titles can be anything, as long as every line carries a name. "The team" is not a name.
The risk
When authority and responsibility are not clearly defined, three things happen:
- Decisions slow down. The minutes when a rollback was still possible are spent finding out who can approve it.Incident calls get crowded. If it is unclear who decides, everyone gets pulled in; a crowded bridge only slows decisions further.The same problem recurs. With no owner for the post-mortem, the root cause is never written up, and a root cause that is not written up is not fixed.
Together these can put you in breach of your SLA commitment with no technical cause at all; time lost to indecision while the system is running is still downtime. A team that never measures how much of that 175-minute average is diagnosis and how much is waiting for approval will optimize the wrong problem.
Eclit note
When we take over an operation, we usually spend the first week filling in these seven lines rather than standing up new monitoring. Technical improvements only pay off once it is written down who makes the decision; otherwise a fault that is detected faster still waits just as long for approval.
We covered the measurement side in our article on observability with Prometheus; this week we look at the human side of the same problem. Last week we verified Kubernetes readiness probes, a check on whether the system reports its own state correctly. Both come down to the same thing: the difference between thinking something works and knowing it does.
Process design and preventive maintenance fall under our system reliability services.
Every Friday. One Production Check.