Beamreach

Guide · Ilia Dubovskii · 10 October 2026

AI Root Cause Analysis for Incidents: What It Can and Can't Do

An AI agent that investigates an alert can save the on-call engineer the worst part of the first half hour: opening ten dashboards, scrolling the deploy log and asking in the channel who changed what. It can't be trusted to decide what's true, or to change production on a hunch. This guide covers where the line falls in practice: what an AI investigation should collect in the first minutes, how it should rank causes, why it should propose rather than act, how to confirm a fix actually worked, how to get post-mortems out of it that people trust, and how to measure the result without fooling yourself.

The short version

1. What AI root cause analysis can and can't do

Incident response has two kinds of work in it. The first is investigation: find out what's affected, what changed, which signals moved and when. It's mechanical, it's parallel, and it's where the first twenty minutes usually go. The second is judgement: decide what's actually going on, whether to roll back, who to wake up, what to tell customers. An AI agent with access to your telemetry can do most of the first and should do very little of the second.

What it does well:

Where it struggles:

This isn't only a cautious vendor's opinion. In August 2025, ClickHouse tested four frontier models on anomalies injected into the OpenTelemetry demo application, with access to the observability data through an MCP server. Their conclusion was blunt: "Autonomous RCA is not there yet." Models improve quickly, so treat any single benchmark as a snapshot. The design conclusion still holds: let the agent find and organise the evidence, and let a human decide.

2. What to collect in the first minutes

The quality of an AI investigation is set by what it can see. Before you worry about the model, make sure the agent can reach these, with read-only access, and that it collects them for every alert:

  1. The alert and the real start time. The alert fires when a threshold is crossed, which is often several minutes after things started going wrong. The agent should find the onset in the metric itself and use that, minus a margin, as the start of every other search window.
  2. What the service depends on, and what depends on it. Take this from APM traces and your service catalog, not a hand-drawn diagram. It tells the agent which databases, queues and downstream services to check, and who else is affected.
  3. Changes before the onset. Deploys, feature flag changes (with who and what percentage), config and secret changes, infrastructure-as-code applies, database migrations, scheduled jobs, autoscaling events. This list finds the cause more often than any anomaly detector.
  4. Dependency health. Database CPU, connections and slow queries; queue depth; cache hit rate; error rates and latency of the services called; third-party status pages.
  5. New log signatures. Not "all errors", but error messages that are new or sharply more frequent than in the baseline window.
  6. Traces of slow or failing requests. Where the time actually goes. This often separates the cause from the symptom.
  7. Similar past incidents. Previous incidents for the same service or with the same signature, and how they were resolved.
  8. Owners and on-call. Who owns each service involved, so the agent can suggest who to page.

The output of this step isn't an answer. It's a set of facts, each with a timestamp and a link back to its source, which is what makes the next two steps checkable.

3. Ranking causes: confidence, evidence and "ruled out"

An investigation that returns one confident sentence is dangerous, because a wrong answer looks exactly like a right one. A useful first evaluation has a structure that lets the on-call engineer check it in a minute:

Treat correlation with suspicion, including when the agent finds it. A change that happened shortly before the onset is a strong lead, not proof. The proof comes later, when you undo the change and the metric recovers.

4. Read-only by default: propose, don't act

The asymmetry decides this. If the agent's hypothesis is wrong, you lose a few minutes reading it. If the agent acts on a wrong hypothesis, it can turn one incident into two: a rollback that removes a fix, a restart that drops in-flight work, a scale-up that hides the real problem and runs up the bill. The vendors in this space mostly agree. Rootly says every change its AI SRE proposes "requires explicit human sign-off before execution". incident.io says the only change its Investigations product can make "is a pull request you review and merge yourself". PagerDuty describes its SRE Agent as performing "approved remediation".

In practice that means:

There's a narrow exception. Some actions are safe enough to run without waiting, if they pass three tests: the outcome is verifiable (you can tell from a metric whether it worked), the action is reversible (one step puts things back), and the blast radius is contained (it touches one component). Turning off a feature flag that was changed ten minutes ago can pass. Failing over a database doesn't. Even then, grant autonomy one action at a time, after the team has watched the agent propose that exact action correctly many times. Our autonomous DevOps framework covers this test and the gate that enforces it in more detail.

5. Confirming recovery on the metric that fired

"Flag turned off" isn't the end of an incident. The end is when the metric that started it is back to normal and stays there. This step is where many incident processes, human or automated, go soft: the action is taken, the channel goes quiet, and nobody checks whether the error rate really came down or just moved somewhere else.

An AI investigation is well placed to close this loop, because it already knows which metric fired and what normal looks like:

This is also what makes the agent improve over time. Each incident ends with a check of whether its top hypothesis was right, which is the number to watch when you decide how much to trust it.

6. Post-mortems you can trust

AI-drafted post-mortems are now a standard feature: incident.io, Rootly, FireHydrant and PagerDuty all offer some form of them. The draft is the easy part. Making it trustworthy takes a few rules:

The quiet benefit is coverage. Small incidents rarely get a write-up because nobody has the time. When the draft costs nothing, more of them get reviewed, and the patterns in small incidents are often what prevents the big ones.

7. Measuring MTTR honestly

Mean time to restore is the obvious metric for an incident agent, and it's the metric we build around, because it's what the business feels. It's also easy to misuse. Two credible sources are worth reading before you put it on a slide:

That doesn't mean you can't measure. It means measuring carefully:

8. Buy or build?

For most teams, buying is the right first move. These are the products we'd look at first, described from their own sites (checked October 2026):

A custom system makes sense in narrower cases: your evidence lives across tools that no single vendor reads well; your runbooks and approval rules are specific and need to be enforced exactly; telemetry can't leave your environment, so the agent must run in your own cloud; or you want the full loop described here (investigate, propose, approve, verify on the metric, draft the post-mortem) tuned to one metric you care about. That's the work we do in Custom Autonomous Systems Engineering.

9. FAQ

How accurate is AI root cause analysis?

It depends heavily on your data and the kind of incident. For failures that follow a recent change or match a known pattern, an agent with good access to telemetry often names the likely cause quickly. For novel or multi-cause incidents it is much weaker: in ClickHouse's August 2025 test of four frontier models, the conclusion was that autonomous root cause analysis is not there yet. Measure it yourself by recording, for every incident, whether the agent's top-ranked cause turned out to be right.

Should an AI incident agent have write access to production?

No, not by default. Give it read-only credentials to observability data, deploy history and the feature flag audit log, and have it propose remediation that a human approves. The only exceptions worth considering are actions that are verifiable, reversible and contained, such as turning off a feature flag that was just changed, granted one action at a time after the team has seen the agent propose it correctly many times.

What data does an AI incident investigation need?

The alert and the metric behind it, service dependencies from APM traces or a service catalog, every change before the onset (deploys, feature flags, config, infrastructure applies, migrations), dependency health, new log signatures, traces of failing requests, similar past incidents, and service owners. Missing any of these creates blind spots, and the agent should say which sources it could not check.

Can AI write the post-mortem?

It can write a good first draft, and most incident management tools now offer this. Build the timeline from timestamped evidence rather than chat, separate what was observed from what was inferred, keep the draft blameless, and present action items as suggestions. A human should review and merge it, for example as a pull request.

Is MTTR a good metric for AI incident response?

It is the metric the business feels, but a mean is misleading because incident durations are heavily skewed. Google's report Incident Metrics in SRE and Courtney Nash's work on the VOID both show how unreliable MTTR can be for trend analysis. Track time to restore per incident, look at the median and a high percentile, and also measure what the agent directly controls: time to first evaluation and how often its top cause was right.

Should we buy an AI SRE tool or build our own?

Buy first if your telemetry mostly lives in one platform and a vendor's workflow fits your on-call process: Datadog, incident.io, Rootly, PagerDuty, FireHydrant and Resolve.ai all offer AI investigation or post-mortem features. Build, or have one built, when your evidence is spread across tools no single vendor reads well, when telemetry can't leave your environment, or when you need your own runbooks and approval rules enforced exactly.

Written by Ilia Dubovskii, founder of Beamreach. Vendor descriptions were checked against each vendor's own site on 10 October 2026. Products in this space change monthly, so if something here no longer matches, the vendor's site wins. Tell us and we'll fix this page.

Related: First Mate demo · Autonomous DevOps · Custom Autonomous Systems Engineering · All guides for cloud teams