The short version
- AI is good at the investigation: gathering evidence from many sources at once and correlating it with recent changes. It is weak at judgement, novel failures and incidents with more than one cause.
- Every hypothesis should come with a confidence, the evidence behind it, and a list of what was checked and ruled out.
- Give the agent read-only credentials. It proposes; a human approves. Only actions that are verifiable, reversible and contained should ever run on their own, one at a time.
- An incident is over when the metric that triggered it is back to normal, not when someone ran a command. The agent should watch that metric and say if the fix didn't work.
- Build post-mortem drafts from timestamped evidence, mark what's inferred, keep them blameless, and have a human merge them.
- Mean time to restore (MTTR) is a noisy number. Look at distributions and per-incident detail, not a single average.
1. What AI root cause analysis can and can't do
Incident response has two kinds of work in it. The first is investigation: find out what's affected, what changed, which signals moved and when. It's mechanical, it's parallel, and it's where the first twenty minutes usually go. The second is judgement: decide what's actually going on, whether to roll back, who to wake up, what to tell customers. An AI agent with access to your telemetry can do most of the first and should do very little of the second.
What it does well:
- Breadth and speed. It can query metrics, logs, traces, the deploy history and the feature flag audit log at once, where a person opens them one tab at a time.
- Correlation with change. Most incidents follow a change: a deploy, a flag, a config edit, a migration, a scaling event. Lining those up against the moment a metric moved is exactly the kind of tedious cross-referencing a model does reliably.
- Recognising known patterns. Connection pool exhaustion, a missing index, an expired certificate, a dependency that's down. If it has happened before, in your systems or in general, the agent will often name it quickly.
- Writing it down. A clear summary in the channel, and later a first draft of the timeline, saves real time.
Where it struggles:
- Novel failures. Without a recognisable pattern or an obvious recent change, a model tends to produce a plausible story rather than say "I don't know".
- Anchoring. Models latch onto the first hypothesis that fits and stop looking, which is a real risk for incidents with more than one contributing cause.
- Missing context. It doesn't know that the marketing campaign started at 9:00, or that the payment provider emailed a maintenance notice, unless that information is somewhere it can read.
- Trade-offs. Whether to roll back a release that also contains a security fix is a business decision, not a pattern match.
This isn't only a cautious vendor's opinion. In August 2025, ClickHouse tested four frontier models on anomalies injected into the OpenTelemetry demo application, with access to the observability data through an MCP server. Their conclusion was blunt: "Autonomous RCA is not there yet." Models improve quickly, so treat any single benchmark as a snapshot. The design conclusion still holds: let the agent find and organise the evidence, and let a human decide.
2. What to collect in the first minutes
The quality of an AI investigation is set by what it can see. Before you worry about the model, make sure the agent can reach these, with read-only access, and that it collects them for every alert:
- The alert and the real start time. The alert fires when a threshold is crossed, which is often several minutes after things started going wrong. The agent should find the onset in the metric itself and use that, minus a margin, as the start of every other search window.
- What the service depends on, and what depends on it. Take this from APM traces and your service catalog, not a hand-drawn diagram. It tells the agent which databases, queues and downstream services to check, and who else is affected.
- Changes before the onset. Deploys, feature flag changes (with who and what percentage), config and secret changes, infrastructure-as-code applies, database migrations, scheduled jobs, autoscaling events. This list finds the cause more often than any anomaly detector.
- Dependency health. Database CPU, connections and slow queries; queue depth; cache hit rate; error rates and latency of the services called; third-party status pages.
- New log signatures. Not "all errors", but error messages that are new or sharply more frequent than in the baseline window.
- Traces of slow or failing requests. Where the time actually goes. This often separates the cause from the symptom.
- Similar past incidents. Previous incidents for the same service or with the same signature, and how they were resolved.
- Owners and on-call. Who owns each service involved, so the agent can suggest who to page.
The output of this step isn't an answer. It's a set of facts, each with a timestamp and a link back to its source, which is what makes the next two steps checkable.
3. Ranking causes: confidence, evidence and "ruled out"
An investigation that returns one confident sentence is dangerous, because a wrong answer looks exactly like a right one. A useful first evaluation has a structure that lets the on-call engineer check it in a minute:
- Ranked hypotheses, each with a plain confidence level (high, medium, low) rather than a falsely precise percentage.
- Evidence for each one, as specific facts with links: "flag rolled from 10% to 100% at 09:41", "database CPU 31% to 97% at 09:43", "93% of slow spans wait on this call". If a hypothesis has no evidence the engineer can click, it shouldn't be ranked first.
- What was ruled out, and why. "No deploys in six hours. Payment provider latency normal. Network metrics at baseline." This is the most underrated part. It stops the team spending fifteen minutes on the reflexive "was it the deploy?", and it shows what the agent actually checked, so a gap is visible.
- Cause versus symptom. An exhausted connection pool is usually a symptom of slow queries, not a cause. The agent should say which it thinks each item is.
- What would change its mind. For example: "if CPU doesn't fall within five minutes of turning the flag off, this hypothesis is wrong".
Treat correlation with suspicion, including when the agent finds it. A change that happened shortly before the onset is a strong lead, not proof. The proof comes later, when you undo the change and the metric recovers.
4. Read-only by default: propose, don't act
The asymmetry decides this. If the agent's hypothesis is wrong, you lose a few minutes reading it. If the agent acts on a wrong hypothesis, it can turn one incident into two: a rollback that removes a fix, a restart that drops in-flight work, a scale-up that hides the real problem and runs up the bill. The vendors in this space mostly agree. Rootly says every change its AI SRE proposes "requires explicit human sign-off before execution". incident.io says the only change its Investigations product can make "is a pull request you review and merge yourself". PagerDuty describes its SRE Agent as performing "approved remediation".
In practice that means:
- Read-only credentials, scoped to observability data, the deploy history and the flag audit log. Not a production admin role "for convenience".
- Remediation as a proposal: a specific action, its expected effect, how to undo it, and a button. A separate, narrowly permissioned path carries out the approved action, and logs who approved it.
- Code changes as pull requests, which go through the same review as any other change.
There's a narrow exception. Some actions are safe enough to run without waiting, if they pass three tests: the outcome is verifiable (you can tell from a metric whether it worked), the action is reversible (one step puts things back), and the blast radius is contained (it touches one component). Turning off a feature flag that was changed ten minutes ago can pass. Failing over a database doesn't. Even then, grant autonomy one action at a time, after the team has watched the agent propose that exact action correctly many times. Our autonomous DevOps framework covers this test and the gate that enforces it in more detail.
5. Confirming recovery on the metric that fired
"Flag turned off" isn't the end of an incident. The end is when the metric that started it is back to normal and stays there. This step is where many incident processes, human or automated, go soft: the action is taken, the channel goes quiet, and nobody checks whether the error rate really came down or just moved somewhere else.
An AI investigation is well placed to close this loop, because it already knows which metric fired and what normal looks like:
- Define "restored" before acting. For example: the error rate below the alert threshold, and close to its baseline, for ten consecutive minutes.
- Watch the cause signals as well as the symptom. If the hypothesis was a slow query, database CPU should fall first, then latency, then errors. If errors drop but CPU doesn't, something else is going on.
- Set a deadline and say so. "If checkout isn't back within ten minutes, I'll report that and reopen the other hypotheses." A fix that didn't work is evidence too.
- Record the times. Onset, alert, first evaluation, action, recovery. These become the post-mortem timeline and the honest version of your MTTR.
This is also what makes the agent improve over time. Each incident ends with a check of whether its top hypothesis was right, which is the number to watch when you decide how much to trust it.
6. Post-mortems you can trust
AI-drafted post-mortems are now a standard feature: incident.io, Rootly, FireHydrant and PagerDuty all offer some form of them. The draft is the easy part. Making it trustworthy takes a few rules:
- Build the timeline from evidence, not from chat. Timestamps should come from the metric, the deploy log and the flag audit trail. Chat messages are useful for what people decided and why, but they record when someone noticed something, not when it happened.
- Separate what was observed from what was inferred. "CPU reached 97% at 09:43" is a fact. "The flag caused the slow query" is a conclusion, confirmed or not by what happened when the flag was turned off. Write it that way.
- Keep it blameless. Google's SRE book puts it plainly: a blameless postmortem must "focus on identifying the contributing causes of the incident without indicting any individual or team". An AI draft can quietly break this by writing "Maria rolled the flag to 100%" under root cause. The change belongs in the timeline. The cause is in the system: a query without an index, and no check of database impact before a flag reaches all traffic.
- Action items are suggestions. Mark them as "for the team to confirm", with suggested owners. Only people should commit people to work.
- Review it like code. Opening the draft as a pull request in the repository where your post-mortems live gives you review, history and an approver by default. incident.io's own docs call its AI draft "a starting point, not a finished document", which is the right framing whatever tool writes it.
The quiet benefit is coverage. Small incidents rarely get a write-up because nobody has the time. When the draft costs nothing, more of them get reviewed, and the patterns in small incidents are often what prevents the big ones.
7. Measuring MTTR honestly
Mean time to restore is the obvious metric for an incident agent, and it's the metric we build around, because it's what the business feels. It's also easy to misuse. Two credible sources are worth reading before you put it on a slide:
- Google's report Incident Metrics in SRE (Štěpán Davidovič, 2021) uses simulation to show that statistics like MTTR are "poorly suited for decision making or trend analysis in the context of production incidents".
- Courtney Nash's work on the VOID, an open database of public incident reports, argues "how unreliable MTTR can be" (SREcon22 Americas). Incident durations are heavily skewed, so a mean says little about a typical incident.
That doesn't mean you can't measure. It means measuring carefully:
- Use the time to restore for each incident, and look at the distribution (median and a high percentile), not a monthly mean.
- Define the start as the onset in the metric, not the alert, and the end as the confirmed recovery from section 5. Otherwise improving the alert threshold looks like improving response.
- Measure the parts the agent controls: time from alert to first evaluation, and how often its top-ranked cause was right. Those are directly attributable. A change in overall MTTR over a few months isn't.
- Compare like with like. A quarter with one long outage will swamp everything else. Read the individual incidents.
- Don't let the number drive behaviour. If people start closing incidents early to keep MTTR down, the metric has stopped measuring anything.
8. Buy or build?
For most teams, buying is the right first move. These are the products we'd look at first, described from their own sites (checked October 2026):
- Datadog Bits AI SRE (called Bits Investigation in Datadog's current docs) investigates alerts as they fire, forms hypotheses from Datadog telemetry and suggests code fixes. The natural choice if Datadog already holds most of your data.
- incident.io has an AI SRE that investigates when an incident is declared, plus AI post-mortem drafts. Its only change to your systems is a pull request you merge.
- Rootly AI SRE ranks evidence-backed hypotheses with confidence scores and suggests fixes, with human sign-off on every change.
- PagerDuty offers an SRE Agent (triage, diagnosis, approved remediation) and a Scribe Agent that captures incident conversations for post-incident reviews.
- FireHydrant offers AI incident summaries and AI-assisted retrospectives. Freshworks has announced it is acquiring FireHydrant.
- Resolve.ai positions itself as an AI SRE that goes on call and mitigates incidents, closer to the autonomous end of the range.
A custom system makes sense in narrower cases: your evidence lives across tools that no single vendor reads well; your runbooks and approval rules are specific and need to be enforced exactly; telemetry can't leave your environment, so the agent must run in your own cloud; or you want the full loop described here (investigate, propose, approve, verify on the metric, draft the post-mortem) tuned to one metric you care about. That's the work we do in Custom Autonomous Systems Engineering.
9. FAQ
How accurate is AI root cause analysis?
It depends heavily on your data and the kind of incident. For failures that follow a recent change or match a known pattern, an agent with good access to telemetry often names the likely cause quickly. For novel or multi-cause incidents it is much weaker: in ClickHouse's August 2025 test of four frontier models, the conclusion was that autonomous root cause analysis is not there yet. Measure it yourself by recording, for every incident, whether the agent's top-ranked cause turned out to be right.
Should an AI incident agent have write access to production?
No, not by default. Give it read-only credentials to observability data, deploy history and the feature flag audit log, and have it propose remediation that a human approves. The only exceptions worth considering are actions that are verifiable, reversible and contained, such as turning off a feature flag that was just changed, granted one action at a time after the team has seen the agent propose it correctly many times.
What data does an AI incident investigation need?
The alert and the metric behind it, service dependencies from APM traces or a service catalog, every change before the onset (deploys, feature flags, config, infrastructure applies, migrations), dependency health, new log signatures, traces of failing requests, similar past incidents, and service owners. Missing any of these creates blind spots, and the agent should say which sources it could not check.
Can AI write the post-mortem?
It can write a good first draft, and most incident management tools now offer this. Build the timeline from timestamped evidence rather than chat, separate what was observed from what was inferred, keep the draft blameless, and present action items as suggestions. A human should review and merge it, for example as a pull request.
Is MTTR a good metric for AI incident response?
It is the metric the business feels, but a mean is misleading because incident durations are heavily skewed. Google's report Incident Metrics in SRE and Courtney Nash's work on the VOID both show how unreliable MTTR can be for trend analysis. Track time to restore per incident, look at the median and a high percentile, and also measure what the agent directly controls: time to first evaluation and how often its top cause was right.
Should we buy an AI SRE tool or build our own?
Buy first if your telemetry mostly lives in one platform and a vendor's workflow fits your on-call process: Datadog, incident.io, Rootly, PagerDuty, FireHydrant and Resolve.ai all offer AI investigation or post-mortem features. Build, or have one built, when your evidence is spread across tools no single vendor reads well, when telemetry can't leave your environment, or when you need your own runbooks and approval rules enforced exactly.
Written by Ilia Dubovskii, founder of Beamreach. Vendor descriptions were checked against each vendor's own site on 10 October 2026. Products in this space change monthly, so if something here no longer matches, the vendor's site wins. Tell us and we'll fix this page.
Related: First Mate demo · Autonomous DevOps · Custom Autonomous Systems Engineering · All guides for cloud teams
Beamreach