Beamreach

Custom autonomous systems · Mock demo

First Mate: an incident companion that investigates every alert

When an alert fires, First Mate works out what's affected, gathers evidence from your monitoring, deploys and feature flags, ranks the likely causes, and posts a first evaluation to the team's channel. After recovery it drafts the post-mortem. It reads production; it never changes it on its own.

The metric at the centre: time to restore service (MTTR). Everything it does is judged by how fast the error rate it is watching comes back down.

Mock demo · fictional company and data

Checkout error rate
0.2%
Normal. The alert fires above 2%.
Target: below 1%

Press Restart, then Next (or →). The run stops when it needs your approval: you play the human.

#checkout-incidents Halyard Outfitters (fictional) in Slack
Thursday · 09:52
📈
Datadogmonitor
[P2] checkout-api: 5xx rate 14% (threshold 2%)
service:checkout-api · env:prod · p99 latency 8.4 s · ALB target 504s rising
⚓
First Mateagent
On it. checkout-api runs on ECS · prod · eu-west-1 and depends on orders-db (Aurora PostgreSQL) and payments-gw. Checking deploys, flags, database and dependencies…
⚓
First Mateagent
SEV-2

Checkout is timing out for about 1 in 7 shoppers

checkout-api · prod · since 09:44 · first evaluation in ~2 min
What's affected

Order placement on the web shop and the mobile app. Browsing, search and the basket are healthy. Payments are not failing; requests time out before they reach the payment step.

Most likely cause, ranked, with evidence
1The recs-at-checkout flag, rolled to 100% at 09:41, added a slow query to every checkoutHigh▸
flagsrecs-at-checkout: 10% → 100% at 09:41 by m.ortiz (LaunchDarkly)
databaseorders-db CPU 31% → 97% at 09:43; top SQL: SELECT … FROM order_items WHERE sku_id = ANY($1) (seq scan, 2.1 s avg)
traces93% of slow checkout spans wait on RecommendationService.fetchAtCheckout()
2Database connection pool exhausted (a symptom of the above)Medium▸
databaseorders-db DatabaseConnections 480/500 since 09:45
logscheckout-api: "timeout acquiring connection from pool" ×1,812 since 09:45
3Ruled out: deploys, payment provider, networkRuled out▸
deploysno ECS deployments to checkout-api or orders-db in the last 6 h
paymentspayments-gw p99 310 ms, error rate 0.1% (normal); provider status page green
networkALB, NAT and VPC flow metrics within baseline
Suggested next steps
  • Turn recs-at-checkout off. Reversible in one click, affects only the recommendation strip.
  • Then add an index on order_items(sku_id) before re-enabling. Draft migration attached as PR #2291.
You're on callOpen PR #2291Page database owner
AI-generated evaluation· a human confirms before anything changes· 7 evidence sources
⚓
First Mateagent
Turned recs-at-checkout off (approved by you, on call). Watching the error rate, database CPU and connection pool. I'll confirm when checkout is back to normal; if it isn't within 10 minutes, I'll say so and look again.
📈
Datadogmonitor
[Recovered] checkout-api: 5xx rate 0.3%
p99 latency 420 ms · orders-db CPU 28% · recovered 5 min after the flag change
⚓
First Mateagent
Recovered in 17 minutes. I've drafted the post-mortem and opened it as a pull request for review.
📝Post-mortem: checkout timeouts, Thursday draft · PR postmortems/checkout-timeouts.mddraft

Summary

For 17 minutes, about 14% of checkout requests timed out. Rolling the recs-at-checkout flag to 100% added an unindexed query to every checkout, saturating the orders database.

Timeline

09:41recs-at-checkout rolled from 10% to 100%. 09:43orders-db CPU reaches 97%; checkout latency climbs. 09:52Datadog alert; First Mate starts the investigation. 09:54First evaluation posted with the flag as the likely cause. 09:56On-call approves; First Mate turns the flag off. 10:01Checkout recovers.

Root cause

The recommendation query filters order_items by sku_id, which has no index. At 10% of traffic the load was absorbed; at 100% it was not.

Action items for the team to confirm

  • Add the index (PR #2291) and load-test before re-enabling. (suggested owner: data platform)
  • Require a database impact check for flags that touch checkout. (suggested owner: checkout team)
  • Alert on orders-db CPU above 80% for 3 minutes.

Drafted by First Mate from the evidence above. A human edits and merges it; nothing is published automatically.

Why it matters

~2 min
from alert to a first evaluation, instead of an engineer starting from a blank screen
17 min
time to restore in this run, with the fix confirmed by the metric, not by a hunch
Every
incident gets a post-mortem draft, so the lessons aren't lost

Figures describe this mock scenario. They are design targets, not measured results.

What it solves

  • The first 20 minutes of every incident. On-call engineers spend them opening dashboards, checking deploys and asking who changed what. First Mate does that in parallel, in about two minutes.
  • Wrong first guesses. "Was it the deploy?" is the reflex. Ranked causes with evidence, including what was ruled out, stop the team chasing the wrong thing.
  • Knowledge that lives in a few heads. Dependencies, owners and past incidents come from your monitoring and ticket history, not from whoever happens to be awake.
  • Post-mortems that never get written. A draft with a timeline built from the evidence turns a chore into a review.

Key features

  • Triggered by your existing alerts: Datadog, CloudWatch, Prometheus, PagerDuty, Opsgenie, uptime checks
  • Knows what depends on what from your APM traces and service catalog, with no infrastructure-mapping project first
  • Evidence from metrics, logs, traces, deploy history, feature flags and recent infrastructure changes
  • Ranked causes with confidence and the evidence behind each, including what it ruled out
  • Suggested remediation as one-click actions that need a human's approval
  • Post-mortem drafts opened as pull requests in your docs repository
  • Lives where your team works: Slack or Microsoft Teams

How it stays safe

First Mate has read-only access to production. Every change it suggests, even a reversible one like turning off a flag, waits for a person to approve it, and after acting it watches the same metric to prove the fix worked. When you trust it with a specific action, you can let that one action run on its own: the operations that are verifiable, reversible and contained, one at a time.