The short version
- Track one number: alerts that reach a human per week, broken down by alert rule. Label each page with what happened next.
- Group by the labels that identify one problem (
group_by), and inhibit symptom alerts while their cause is firing (inhibit_ruleswith a correctequallist). - Page only for
severity="page". Everything else goes to a ticket queue or a channel, never to a phone. - Stop flapping with
for:andkeep_firing_for:, and test rule changes withpromtool test rules. - Automate a runbook only when its result can be checked, undone and kept small.
- Review the noisiest rules weekly and change them by pull request. After every change, check that real incidents would still have paged.
1. Measure the noise before you touch a rule
"We have too many alerts" is a feeling. To fix it you need a number that goes down when things improve and doesn't go down when you simply break alerting. The one that works: alerts that reach a human per week. That means pages, phone calls and messages that someone is expected to look at, not everything Prometheus evaluates.
There are three places to get it, from easiest to most complete:
- Your paging tool. PagerDuty, Opsgenie and Grafana IRM all keep a history of incidents or alert groups you can pull from the UI or the API. Count them per week and group them by alert name.
- Alertmanager's own metrics.
alertmanager_notifications_totalcounts notifications sent, labelled byintegration(for examplepagerdutyorslack).sum by (integration) (increase(alertmanager_notifications_total[7d]))gives a weekly total per channel. It counts notifications, not alert groups, so repeats and resolved messages are included. Use it for the trend, not as an exact count. - A webhook log. Add a
webhook_configsreceiver that writes every notification payload to a file or table. The payload includes each alert's labels,startsAtandfingerprint, so you can count per rule and per firing.
To see how much time each rule spends firing, Prometheus already has the data. It writes a synthetic ALERTS series with alertstate="firing" at every rule evaluation, so the number of samples is the time spent firing divided by the evaluation interval:
# Samples spent firing per rule over the last week.
# Multiply by your rule evaluation interval to get firing time.
topk(20, sum by (alertname) (count_over_time(ALERTS{alertstate="firing"}[7d])))
Then add the part no tool gives you: what happened after each page. For a few weeks, have whoever is on call tag every page with one of four outcomes: acted (did something that changed the system), looked, nothing to do, duplicate (another page already covered it) or already known. A rule's actionability rate is acted divided by total pages. It's a rough measure, and that's fine. A rule that paged 40 times last month with 2 actions is obvious without decimals.
2. Sort every paging rule: page, ticket, automate or delete
The Google SRE book's chapter on monitoring has the best short checklist for this. For each rule that can page, ask whether it detects a condition that is urgent, actionable and user-visible that nothing else detects; whether you'll ever ignore it because it's benign; whether the action could wait until morning or be safely automated; and whether other people are being paged for the same issue. The chapter puts it in one line: "Every page should be actionable."
The answers put each rule into one of four buckets:
| Bucket | Test | Where it goes |
|---|---|---|
| Page | Users are affected or soon will be, and a person must act now | On-call, with a runbook link |
| Ticket | Needs fixing, but tomorrow is fine (a certificate expiring in 14 days, a disk at 70%) | Ticket queue, deduplicated |
| Automate | The response is always the same, and its result can be checked | A runbook that runs and verifies, paging only if it fails |
| Delete | Nobody has acted on it in months, or another alert always fires with it | Removed, after the check in section 9 |
Prefer paging on symptoms users feel (error rate, latency, failed checkouts) over causes (CPU, memory, a single pod restarting). Cause alerts are useful as context and as tickets. As pages, they're how one bad deploy turns into 30 notifications.
3. Group alerts that describe one problem
Alertmanager batches alerts with the same values for the labels in group_by into one notification. Many configs group too finely, or use group_by: ['...'], which the docs say "effectively disables aggregation entirely". Pick the labels that identify one problem from the responder's point of view, usually cluster, namespace or service, plus the alert name:
route:
receiver: ticket-queue # anything not matched below never pages
group_by: [cluster, namespace, alertname]
group_wait: 30s # wait for the rest of the group before the first notification
group_interval: 5m # how often to send updates about a group that changed
repeat_interval: 4h # re-send an unchanged, still-firing group
routes:
- receiver: oncall-pager
matchers:
- severity="page"
- receiver: team-channel
matchers:
- severity="info"
Three details the docs spell out and teams tend to miss:
group_waitalso decides whether inhibition works. If it's too short, the first notification goes out before the alert that should inhibit it has arrived. 30 seconds (the default) is a reasonable floor when you rely on inhibition.repeat_intervalis rounded up to a multiple ofgroup_interval, so set it as one.- Child routes inherit
group_by,group_wait,group_intervalandrepeat_intervalfrom the parent unless they set their own. A team route can group byservicewhile the root groups bycluster.
Grafana Alerting works the same way: notification policies have a Group by option and the same three timers, and by default group by alertname and grafana_folder.
4. Inhibit symptoms while the cause is firing
Grouping merges alerts that look alike. Inhibition handles alerts that look different but share a cause: a node runs out of disk, and you also get evictions, crash loops and 5xx errors from everything scheduled on it. An inhibit rule mutes the target alerts while a source alert is firing, as long as the labels in equal match:
inhibit_rules:
# A critical alert hides the warning version of the same alert.
- source_matchers:
- severity="critical"
target_matchers:
- severity="warning"
equal: [alertname, cluster, service]
# Disk pressure on a node hides pod symptoms on that same node.
# Alert names are examples: use the names your rules actually produce.
- source_matchers:
- alertname="NodeDiskPressure"
target_matchers:
- alertname=~"KubePodEvicted|KubePodCrashLooping"
equal: [cluster, node]
The second rule is where most inhibition goes wrong, so check two things.
Both alerts must carry every label in equal. Many pod-level alerts don't have a node label, because the metric they're built on doesn't. If the source has node="ip-10-2-14-7" and the target has no node at all, the values differ and nothing is inhibited. The fix is to add the label in the rule's expression (for example by joining with kube_pod_info), not to drop node from equal.
Missing labels count as equal. The Alertmanager docs say: "if all the label names listed in equal are missing from both the source and target alerts, the inhibition rule will apply." An equal list with a typo in a label name can therefore mute every matching target in every cluster. Keep the source and target matchers narrow, and test the rule (section 9).
Inhibited alerts aren't deleted. They still show in Prometheus and in the Alertmanager API as suppressed, with the inhibiting alert listed. That matters when you check you lost nothing.
5. Route non-urgent alerts to tickets, not phones
The cheapest change on this list: make paging opt-in. With the routing tree above, only severity="page" reaches the on-call receiver. Everything else lands in a ticket queue through a webhook:
receivers:
- name: oncall-pager
pagerduty_configs:
- routing_key_file: /etc/alertmanager/secrets/pagerduty-key
- name: ticket-queue
webhook_configs:
- url: https://alert-to-ticket.internal.example/hook
send_resolved: true
- name: team-channel
slack_configs:
- api_url_file: /etc/alertmanager/secrets/slack-webhook
channel: '#ops-alerts'
The webhook service should create one ticket per alert group (use the payload's groupKey), comment on it on repeats instead of opening a new one, and close it when the alert resolves. That turns the cert-expiring-in-14-days alert from six pages a week into one ticket.
Check routing changes without deploying them: amtool check-config alertmanager.yml validates the file, and amtool config routes test --config.file=alertmanager.yml severity=page service=payments prints which receiver a given label set would reach.
6. Tune thresholds with for: and keep_firing_for:
A rule that fires for 40 seconds and resolves on its own is flapping. Two fields fix most of it:
groups:
- name: payments
rules:
- alert: PaymentsHighErrorRate
expr: |
sum(rate(http_requests_total{job="payments", code=~"5.."}[5m]))
/ sum(rate(http_requests_total{job="payments"}[5m])) > 0.02
for: 10m # must be true for 10 minutes before it fires
keep_firing_for: 5m # stays firing 5 minutes after it stops being true
labels:
severity: page
annotations:
summary: Payments 5xx above 2% for 10 minutes
runbook_url: https://runbooks.internal.example/payments-5xx
for keeps the alert in the pending state until the condition has held for that long, so short spikes never notify. keep_firing_for (Prometheus 2.42 and later) keeps it firing for a while after the condition clears, so a metric that dips below the threshold for one evaluation doesn't produce a resolve and a fresh page. Grafana alert rules have the same idea as a pending period.
Don't guess the durations. Use the firing-time query from section 1 and your incident history: if real incidents for this rule always lasted longer than 15 minutes and the noise always cleared in under 3, a for of 5 to 10 minutes costs you little detection time. Write the decision down as a unit test so the next person can see why:
# payments_test.yml, run with: promtool test rules payments_test.yml
rule_files:
- payments_rules.yml
evaluation_interval: 1m
tests:
- interval: 1m
input_series:
# 3-minute spike to 5% errors, then normal: must NOT fire.
- series: 'http_requests_total{job="payments", code="500"}'
values: '0+0x10 0+5x3 15+0x20'
- series: 'http_requests_total{job="payments", code="200"}'
values: '0+100x40'
alert_rule_test:
- eval_time: 20m
alertname: PaymentsHighErrorRate
exp_alerts: []
Add a second test with a sustained error rate that must fire. The pair stops anyone loosening the rule until it misses a real outage.
7. Which runbooks are safe to automate
If the response to an alert is always the same (clear old logs, restart one stuck worker, rotate one node), a person doing it at 2 a.m. adds delay and little judgement. But automatic remediation is only safe when three things are true. We call them verifiable, reversible and contained:
- Verifiable. You can tell from a metric whether it worked, ideally the same metric that raised the alert. "Disk on the node below 80% within 5 minutes" is verifiable. "Restarted the service" is not; it only says the command ran.
- Reversible. If it was the wrong move, you can undo it. Restarting a stateless pod is reversible in practice. Deleting data or failing over a primary database usually isn't.
- Contained. It acts on a bounded scope with hard limits: one node at a time, at most three runs an hour, never during a deploy, never on more than 10% of a pool.
Wire it so the runbook runs, waits, reads the triggering metric again, and if the metric hasn't recovered, stops and pages a person with what it tried. An automated fix that fails quietly is worse than a page. Log every run where the team will see it, so the weekly review can count automated fixes too: a rule fixed automatically 30 times a week is a root cause waiting for a ticket.
8. The weekly noise review
Alert rules decay as the system under them changes, so tuning once doesn't hold. A 30-minute weekly slot with the people who were on call does. A format that works:
- The number (5 minutes). Alerts that reached a human this week versus last, and how many were acted on.
- The top five rules by pages (20 minutes). For each: how many pages, how many acted on, how long it fired. Decide one of: keep, retune, demote to ticket, automate, delete. Write down why.
- Owners (5 minutes). Every decision becomes a pull request against the rules repo with a named owner, merged by the following week.
Two rules keep it honest. Changes go through pull requests, never through a silence that someone forgets to remove. And nobody mutes a rule because it's annoying without saying what would catch the problem instead.
9. Check you didn't lose signal
You can always reduce alerts by alerting on less. The number from section 1 can't tell you whether you did that, so every change needs a second check.
- Replay real incidents. Take the last few months of real incidents, especially those found by a page. For each, ask whether the changed rules would still have paged, and how much later. Unit tests make this concrete: feed the incident's metric shape into
promtool test rulesand assert the alert still fires. - Watch suppressed alerts for a week. Prometheus's
ALERTSseries isn't affected by Alertmanager inhibition or routing. Compare what fired in Prometheus with what was delivered. Any alert that was inhibited while its source alert was not about the same problem is a badequallist. - Count incidents found the wrong way. Track incidents first reported by customers, support or another team. If that count rises after a tuning round, you cut too deep, whatever the page count says.
- Demote before deleting. Move a doubtful rule to
severity="ticket"for a month first. If nobody acts on any of those tickets either, delete it.
10. Where products and agents fit
Everything above works with open-source Prometheus and Alertmanager, and for many teams it's enough. The commercial tools automate parts of it:
- PagerDuty offers Intelligent Alert Grouping (machine-learned grouping of related alerts into one incident) and Auto-Pause Incident Notifications (holding notifications for transient alerts so they can resolve on their own) in its AIOps add-on, or as Signal Intelligence in its PD Reliability Platform plans. PagerDuty Runbook Automation can run diagnostics and remediation from incidents.
- Datadog Event Management groups related alerts into cases, with Intelligent Correlation (machine learning using service topology) and Pattern-based Correlation (rules you define).
- BigPanda correlates alerts from many monitoring tools into incidents and is aimed at large enterprise IT operations teams.
- Opsgenie stopped being sold on 4 June 2025 and its support ends on 5 April 2027. Atlassian moves its alerting and on-call features into Jira Service Management, so plan that migration rather than new Opsgenie tuning.
What none of them do for you is the judgement in sections 2, 7, 8 and 9: deciding which rule is wrong, fixing it in your repo, and proving nothing was lost. That's the part an AI agent built for your environment can take on. Lookout is a mock demo of that design. It groups alerts by shared labels, timing and cause, runs only pre-approved runbooks with limits and checks each fix against the triggering metric, attaches alerts to existing tickets, and runs this guide's weekly review itself, proposing rule changes as pull requests and never silencing a rule on its own. It's built around one number, alerts that reach a human. See this design working, or read how we approach custom autonomous systems.
11. FAQ
What is alert fatigue?
Alert fatigue is what happens when on-call engineers get so many alerts that lead nowhere that they stop treating each one as urgent. Pages get acknowledged without being read, and the alert that matters is missed or answered late. The fix is fewer alerts reaching people, not faster responses to all of them.
How do I reduce the number of alerts from Alertmanager?
Start by counting pages per alert rule for a few weeks and noting which ones anyone acted on. Then group alerts that describe one problem with group_by, inhibit symptom alerts while their cause is firing with inhibit_rules, route anything non-urgent to a ticket queue instead of the pager, and add for: durations to rules that flap. Review the noisiest rules weekly and change them by pull request.
What is the difference between grouping and inhibition in Alertmanager?
Grouping merges alerts that have the same values for the labels in group_by into one notification, for example nine crash-looping pods in one namespace. Inhibition mutes alerts of a different kind while a related source alert is firing, for example pod evictions on a node that is out of disk. Grouping reduces repeats; inhibition removes symptoms of a cause you already know about.
Why is my Alertmanager inhibit rule not working?
The usual cause is the equal list. Both the source and target alerts must have the same value for every label in it, and if the target alert doesn't carry a label such as node at all, the values differ and nothing is inhibited. A group_wait that's too short is the other cause: the target's notification goes out before the source alert arrives. Check the labels on both alerts in the Alertmanager UI or with amtool.
How long should the for: duration be on a Prometheus alert?
Long enough that short spikes which clear on their own never fire, and short enough that a real incident still pages in time. Look at how long this rule's noisy firings lasted compared with real incidents, and pick a value between them. Add keep_firing_for (Prometheus 2.42 and later) if the alert resolves and re-fires when the metric hovers around the threshold.
Which runbooks are safe to automate?
Runbooks whose result you can check with a metric, that you can undo, and that act on a small, bounded scope, such as clearing old logs on one node or restarting one stateless worker. The automation should re-read the metric that triggered the alert after it runs and page a person if the metric hasn't recovered. Anything that deletes data or fails over a database should stay with a human.
How do I know I didn't lose important alerts after tuning?
Replay recent real incidents against the changed rules, ideally as promtool unit tests, and confirm each one would still have paged. Compare what fired in Prometheus with what Alertmanager delivered for a week to catch inhibit rules that hide too much. Track incidents first reported by customers or other teams: if that number rises after a tuning round, you cut too deep.
Written by Ilia Dubovskii, founder of Beamreach. Config was checked against the Prometheus and Alertmanager documentation, and product details against each vendor's own site, on 10 October 2026. If something here no longer matches the docs, the docs win. Tell us and we'll fix this page.
Related: Lookout: AI alert triage demo · Autonomous DevOps · Custom autonomous systems · All guides for cloud teams
Beamreach