Beamreach

Custom autonomous systems · Mock demo

Cut alert fatigue: Lookout pages a person only when an alert needs one

Lookout sits between your monitoring and your on-call rota. It groups alerts that share a cause, runs the runbooks you've approved for routine problems, and sends a person only what needs one, with the evidence attached. Every Friday it shows which alert rules create the most noise and proposes how to fix them.

The metric at the centre: alerts that reach a human. Lookout's job is to make that number small and keep it honest.

Doing this by hand first? Our guide to reducing alert fatigue in Prometheus and Alertmanager covers grouping, inhibition, routing, for: durations and the weekly noise review.

Mock demo · fictional company and data

Alerts sent to a human this week
486
Last week, before Lookout: 486 of 512 alerts paged someone.
Target: under 50 a week

Press Restart, then Next (or →). The run stops when it needs your approval: you play the human.

#ops-alerts Cobalt Payments (fictional) in Slack
Monday · 02:14
🔔
Alertmanagermonitor
37 alerts firing in 6 minutes
NodeDiskPressure ×3 · PodEvicted ×14 · KubePodCrashLooping ×9 · HTTP5xxHigh ×3 · CertExpiringSoon ×6 · PostgresReplicaLag ×2
🔭
Lookoutagent
37 → 3

37 alerts, 3 actual problems

grouped by shared labels, timing and cause · in 50 s
Problems, with how each will be handled
ADisk full on 3 nodes; the evictions, crash loops and 5xx are symptoms (29 alerts)Runbook▸
nodesip-10-2-14-7, -15-2, -15-9: /var/log 97–99% full
logspayments-worker writing 1.8 GB/h since Friday's deploy: LOG_LEVEL=debug left on
policyrunbook RB-12 "node disk pressure" approved to run on its own, one node at a time
BStaging certificates expire in 9 days (6 alerts)Known▸
jiraOPS-1180 "renew staging certs" open, due Thursday
actionattach to OPS-1180, silence until Thursday 18:00, re-alert if not closed
Cpayments-db replica 45 s behind and growing (2 alerts)Human▸
postgresreplica lag 12 s → 45 s in 20 min; no runbook covers this
activitypg_stat_activity: query by bi_reader running 38 min on the replica, holding replay
Grouping is a draft until verified· if a fix doesn't clear its group, Lookout re-checks and escalates
🔭
Lookoutagent
Problem A fixed with runbook RB-12. One node at a time: cordoned, rotated logs over 1 GB, uncordoned. Verified: DiskPressure=False on all 3 nodes, evicted pods rescheduled, 5xx back to 0.1%. 29 alerts resolved. Opened PAY-2231 for the payments team: LOG_LEVEL=debug is still on and the disks will fill again in about 30 hours.
🔭
Lookoutagent
Paged @payments-oncall for problem C, with the evidence above. Likely fix: cancel the 38-minute bi_reader query on the replica, but it belongs to a finance report, so that's a person's call. Problem B is attached to OPS-1180 and silenced until Thursday.
✅ ack by @r.santos 02:36
Friday · 16:00
🔭
Lookoutagent
Weekly review

3 alert rules caused 371 of this week's 497 alerts, and none led to a human action

last 4 weeks of alert history and on-call actions · proposal: PR #77 in alert-rules
Noisiest rules and the proposed change
disk_usage_80212 fires in 4 weeks, 0 actions. Raise to 90% for 10 min; RB-12 handles the rest.Tune
pod_restart_any96 fires, every one a rolling deploy. Alert only on more than 3 restarts in 10 min.Tune
latency_p99_staging63 fires, staging only. Send to the ticket queue instead of the pager.Reroute
You're the platform leadReview the diff
Changes to alerting always need a person· Lookout never silences a rule on its own
🔭
Lookoutagent
PR #77 merged and the new rules are live. I'll watch for anything the old rules would have caught that the new ones miss, and report the real numbers next Friday.
Next Friday · 16:00
🔭
Lookoutagent
Weekly review: 162 alerts (down from 497), 19 reached a human (down from 41). Checked: no incident this week started with a signal the old rules would have caught first. PAY-2231 is fixed, so RB-12 didn't need to run.
📊On-call report: week 2 sent to #platform-leads · stored in ops-reports/

Alerts that reached a human

Before Lookout486 of 512 Week 141 of 497 Week 219 of 162

Handled without a page

  • Runbooks run: 4 (RB-12 ×3, RB-04 ×1), all verified.
  • Grouped as symptoms of another alert: 88.
  • Attached to existing tickets: 21.

Suggested next

  • Write a runbook for replica lag caused by long queries (paged twice this month).

Why it matters

486 → 19
alerts a week that need a person, in this run
1 page
at 2 a.m. instead of 37, with the evidence already gathered
Every
runbook run checked against the metric that triggered it

This is a mock scenario. The figures are design targets, not measured results.

What it solves

  • Alert fatigue. When most pages lead nowhere, people stop reading them, and the real one gets missed.
  • One cause, dozens of pages. A full disk shows up as evictions, crash loops and errors. Lookout pages once, for the cause.
  • Routine fixes done by hand at 2 a.m. The problems your runbooks already cover get fixed and checked without waking anyone.
  • Noisy rules nobody has time to fix. The weekly review turns alert history into specific rule changes, measured afterwards.

Key features

  • Takes alerts from Alertmanager, Datadog, CloudWatch, Grafana, PagerDuty or Opsgenie
  • Groups alerts by shared labels, timing and cause, and pages once per problem
  • Runs only the runbooks you've approved to run on their own, with limits (such as one node at a time)
  • Checks every fix against the metric that raised the alert
  • Escalates with the evidence and a suggested next step
  • Links alerts to existing tickets instead of paging again
  • Weekly noise review with proposed rule changes as pull requests

Buy or build?

Try what you already have first. Most teams can cut a lot of noise with the tools below, and some never need more. Product details were checked on each vendor's site in October 2026.

Alertmanager or Grafana Alerting, configured well

Free and already running for most Prometheus and Grafana users. Grouping (group_by), inhibit rules, routing non-urgent alerts to tickets and for: durations cut out repeat and symptom pages. Enough when you have one monitoring stack, consistent labels and someone with time to review rules. Our step-by-step guide covers it.

PagerDuty AIOps and Runbook Automation

Intelligent Alert Grouping uses machine learning to merge related alerts into one incident, and Auto-Pause holds notifications for transient alerts so they can resolve on their own. Both come with the AIOps add-on, or as Signal Intelligence in PD Reliability Platform plans. Runbook Automation can run diagnostics and fixes from an incident. Enough when PagerDuty is already where every alert ends up.

Datadog Event Management

Groups related alerts into cases, with Intelligent Correlation (machine learning that uses your service topology) or Pattern-based Correlation (rules you define). Enough when most of your monitoring already lives in Datadog.

BigPanda

An "agentic ITOps" platform (its own term) that correlates alerts from many monitoring and observability tools into incidents, with ServiceNow and Jira integrations. Enough when you're a large IT operations organisation with many tools and an ITSM process to feed.

On Opsgenie? Atlassian stopped selling it on 4 June 2025 and ends support on 5 April 2027; its alerting and on-call features move to Jira Service Management. Plan that migration before investing in Opsgenie-specific tuning.

When a custom build makes sense

  • Alerts come from several tools (Alertmanager, CloudWatch, Datadog) and no single product sees all of them.
  • You want routine fixes to run unattended, but only your approved runbooks, with your limits, each checked against the metric that raised the alert.
  • You want noisy rules fixed in your own repo, as pull requests your team reviews, not tuned inside a vendor's console.
  • You want one number, alerts that reach a human, reported every week with the evidence behind it.

That's what we build. See Custom Autonomous Systems Engineering.

Questions

How can AI reduce alert fatigue for on-call engineers?

By deciding what reaches a person. An AI triage agent groups alerts that share a cause, handles the routine ones with runbooks you've approved, and pages someone only for what's left, with the evidence already collected. It also reviews which alert rules create the most noise and proposes fixes, so the number of pages keeps falling.

What is alert correlation, and how is it different from Alertmanager grouping?

Alertmanager grouping batches alerts that have the same values for the labels you choose, such as cluster and alertname. Correlation goes further and links alerts that look different but share a cause, for example a full disk and the pod evictions and 5xx errors it causes. Alertmanager's inhibit rules handle simple cause-and-symptom pairs (how to set them up); correlation tools and agents use timing, topology and history to handle the rest.

Can an AI agent run runbooks automatically without making things worse?

Only within limits. Lookout runs a runbook on its own only if your team has approved it for that, and only within the scope you set, such as one node at a time. After each run it checks the metric that raised the alert; if the metric hasn't recovered, it stops and pages a person with what it tried. Runbooks that can't be checked, undone or kept small stay with humans.

Will an AI triage agent silence alerts I need?

Lookout never silences or changes an alert rule on its own. Its weekly noise review proposes rule changes as pull requests that your team reviews and merges. After a change, it compares what fired with what was delivered, so a rule that went quiet when it shouldn't have shows up in the next review.

How do I measure alert noise?

Count the alerts that reach a human each week, broken down by alert rule, and note for each page whether anyone acted on it. Rules that page often and rarely lead to action are your noise. Lookout is built around that single number, alerts that reach a human, and reports it weekly. The guide shows where to get the numbers in Prometheus and Alertmanager.

Does Lookout replace PagerDuty or Alertmanager?

No. It sits between them. Alertmanager, Datadog, CloudWatch or Grafana still detect problems, and PagerDuty or your on-call tool still pages people. Lookout decides what gets passed on, what gets fixed by an approved runbook, and what gets attached to an existing ticket instead.

How it stays safe

Lookout acts on its own only through runbooks your team has approved for that, within the limits you set, and each one is verifiable, reversible and contained. If a fix doesn't move the metric, it stops and pages a person. It never silences or changes an alert rule without approval, and after a change it checks that nothing important went quiet.