Your MTTR Will Improve This Year Whether or Not You Got Better at Incidents

TV
Thiago Victorino
6 min read
Your MTTR Will Improve This Year Whether or Not You Got Better at Incidents

Mean time to resolution is an average, and an average is a claim about a population. Change which incidents fall into that population and the number moves without anyone getting better at anything.

That is what AI incident responders are doing to MTTR right now. They absorb the disk-full alert, the expired certificate, the rollback whose culprit is the last deploy. Those were always the bulk of the count and the bottom of the difficulty distribution. Pull them out of the human queue, close them without a human, and the mean drops. The organization’s ability to survive a genuinely novel failure is untouched by that arithmetic. The dashboard will suggest otherwise, and it will do so in green.

The Prediction, and Who Is Making It

Sylvain Kalache, who leads AI Labs and developer relations at Rootly, states the split plainly: “the average MTTR for most incidents will go down…but…resolution time will shoot up for complex incidents.”

Two things about that sentence before we build on it. First, Rootly sells incident simulation, which is the remedy Kalache’s article recommends. Second, the prediction is his own and it is unmeasured. Nobody has published the segmented series that would confirm or kill it.

So the prediction is a hypothesis. What makes it worth the reader’s time is that the mechanism underneath it is forty years old, well documented, and independently checkable.

The 1983 Version of This Argument

Lisanne Bainbridge, a human-factors researcher, wrote Ironies of Automation in 1983. The line Kalache anchors on is the whole argument in one clause: automation “reduces operators’ opportunities to practice routine work while leaving them responsible for new and abnormal situations.”

Read the two halves as a ledger. On one side, reps removed. On the other, responsibility retained. The routine work that automation takes away is also the work that builds the mental model an engineer uses when something unprecedented breaks. You do not learn how a distributed system behaves under stress by reading the architecture diagram. You learn it by being paged at 3am through a long run of boring, mostly-solvable failures and slowly forming an intuition about which subsystem lies to you.

Kalache’s formulation of the consequence is the sentence worth pinning above the incident channel: “the more successful it becomes, the less prepared humans may be for the moment it fails.”

Note the hedge in his own words. May be. He is not claiming to have measured degradation. Neither am I.

Aviation Priced This Trade a Long Time Ago

The comparison Kalache reaches for is jet engines, and it is the useful one because the numbers exist.

Modern commercial aviation runs at fewer than one in-flight shutdown per 100,000 engine flight hours. That reliability is real, earned, and the reason nobody thinks twice about boarding. It also means a working pilot can go an entire career without handling the failure they are certified to handle.

When the failure does arrive, the clock is short. TransAsia Airways Flight 235 crashed 117 seconds after the first warning. Whatever the crew was going to bring to that problem, they had already brought it before the first warning fired. There was no time to acquire it.

Aviation’s answer to that arithmetic is a mandated cadence: FAA rules require captains to complete recurrent training or a proficiency check every six months. What I take from that rule is its shape rather than its reasoning: it treats practice as something scheduled in advance, ahead of any observed decay.

I am not going to tell you six months is the right interval for software incident response. That number was set for a different failure domain by people with accident data I do not have. The transferable part is the shape of the policy: practice is scheduled on a cadence, it is mandatory, and it is decoupled from whether anything has broken recently.

The Fix Is a Metric Split

Rolling back AI incident response is not on the table and should not be. It works. The cheap and immediate correction is to stop reporting one number.

Segment MTTR into buckets before it reaches a dashboard. A workable starting cut:

  • Incidents closed by automation with no human in the loop.
  • Incidents closed by a human with AI assistance.
  • Incidents that required human diagnosis, where the AI contributed context but not the resolution.

Then report each series separately, and report the mix alongside them. The mix is the part worth reporting first, because it is what moves the aggregate. If the routine bucket grows from a minority of incident volume to the clear majority, the aggregate MTTR will fall for that reason alone, and the third bucket can be getting slower the entire time without the top-line number blinking.

That third series is your actual operational health. It is the small-N, high-variance, embarrassing one. It is also the only one that tells you what happens on the day the automation is the thing that is wrong.

Two guardrails on the split. Assign the complexity class at resolution time rather than at page time, because severity labels set during an incident encode panic more than difficulty. And keep the buckets stable for at least a year, because a taxonomy that gets redefined every quarter produces trend lines that mean nothing.

The Same Shape, Three Other Places

An aggregate that improves while its tail degrades is a general failure of measurement, and incident response is just where it is currently most visible.

Token price is the clearest sibling. The per-token price is not the unit that carries retries, re-verification and abandoned attempts, so it can fall while the cost of a finished task does not. We argued the accounting version of this in cost per completed task: the unit that got cheaper is not the unit the business buys. The oversight labor that makes AI output usable shows up in the same place, which is a cost line worth drawing.

The third instance is a green test run. Coverage percentage is an aggregate over assertions, and it rises happily while the assertions that would catch a real regression go unwritten. Same structure. A number improves because its easy population grew.

And when AI is doing the verifying of AI in incident response, the loop itself needs a check, which is the argument in the re-verification loop. That is a mechanism question. This one is a measurement question, and the measurement question is cheaper to fix.

Do This Now

Open your incident tracker. Pull the last two quarters. Tag each incident with one of the three buckets above, by hand if you have to, and plot MTTR per bucket rather than in total. This is a short piece of manual work.

You are looking for one thing: whether the human-diagnosis series is flat, falling, or rising while the aggregate falls. If it is rising, you have a number to bring to the next operational review, and it is a number nobody in the room currently has.

If the three series turn out to be indistinguishable, you have learned something too, and you have built the instrument you will need when they diverge. Which they will, because the mix is going to keep shifting whether or not you are measuring it.


This analysis synthesizes AI handles incidents, engineers lose touch with their systems (Sylvain Kalache, AI Labs lead and DevRel at Rootly, September 2026).

Victorino Group helps engineering organizations instrument AI-assisted operations so the metrics reported to leadership survive contact with the incidents that matter. Let’s talk.

All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →

If this resonates, let's talk

We help companies implement AI without losing control.

Schedule a Conversation