Beamreach

Guide · Ilia Dubovskii · 10 October 2026

How to Reduce Alert Fatigue in Prometheus and Alertmanager: Measure, Cut, Verify

Most advice on alert fatigue stops at "make alerts actionable". This guide is the working version for teams on Prometheus and Alertmanager: how to count the alerts that actually reach a person, which config changes cut them (grouping, inhibition, routing, for: durations), which runbooks are safe to automate, how to run a weekly noise review, and how to check afterwards that you didn't throw away the alert that mattered. Every config snippet was checked against the Prometheus and Alertmanager documentation.

The short version

1. Measure the noise before you touch a rule

"We have too many alerts" is a feeling. To fix it you need a number that goes down when things improve and doesn't go down when you simply break alerting. The one that works: alerts that reach a human per week. That means pages, phone calls and messages that someone is expected to look at, not everything Prometheus evaluates.

There are three places to get it, from easiest to most complete:

To see how much time each rule spends firing, Prometheus already has the data. It writes a synthetic ALERTS series with alertstate="firing" at every rule evaluation, so the number of samples is the time spent firing divided by the evaluation interval:

# Samples spent firing per rule over the last week.
# Multiply by your rule evaluation interval to get firing time.
topk(20, sum by (alertname) (count_over_time(ALERTS{alertstate="firing"}[7d])))

Then add the part no tool gives you: what happened after each page. For a few weeks, have whoever is on call tag every page with one of four outcomes: acted (did something that changed the system), looked, nothing to do, duplicate (another page already covered it) or already known. A rule's actionability rate is acted divided by total pages. It's a rough measure, and that's fine. A rule that paged 40 times last month with 2 actions is obvious without decimals.

2. Sort every paging rule: page, ticket, automate or delete

The Google SRE book's chapter on monitoring has the best short checklist for this. For each rule that can page, ask whether it detects a condition that is urgent, actionable and user-visible that nothing else detects; whether you'll ever ignore it because it's benign; whether the action could wait until morning or be safely automated; and whether other people are being paged for the same issue. The chapter puts it in one line: "Every page should be actionable."

The answers put each rule into one of four buckets:

BucketTestWhere it goes
PageUsers are affected or soon will be, and a person must act nowOn-call, with a runbook link
TicketNeeds fixing, but tomorrow is fine (a certificate expiring in 14 days, a disk at 70%)Ticket queue, deduplicated
AutomateThe response is always the same, and its result can be checkedA runbook that runs and verifies, paging only if it fails
DeleteNobody has acted on it in months, or another alert always fires with itRemoved, after the check in section 9

Prefer paging on symptoms users feel (error rate, latency, failed checkouts) over causes (CPU, memory, a single pod restarting). Cause alerts are useful as context and as tickets. As pages, they're how one bad deploy turns into 30 notifications.

3. Group alerts that describe one problem

Alertmanager batches alerts with the same values for the labels in group_by into one notification. Many configs group too finely, or use group_by: ['...'], which the docs say "effectively disables aggregation entirely". Pick the labels that identify one problem from the responder's point of view, usually cluster, namespace or service, plus the alert name:

route:
  receiver: ticket-queue          # anything not matched below never pages
  group_by: [cluster, namespace, alertname]
  group_wait: 30s                 # wait for the rest of the group before the first notification
  group_interval: 5m              # how often to send updates about a group that changed
  repeat_interval: 4h             # re-send an unchanged, still-firing group
  routes:
    - receiver: oncall-pager
      matchers:
        - severity="page"
    - receiver: team-channel
      matchers:
        - severity="info"

Three details the docs spell out and teams tend to miss:

Grafana Alerting works the same way: notification policies have a Group by option and the same three timers, and by default group by alertname and grafana_folder.

4. Inhibit symptoms while the cause is firing

Grouping merges alerts that look alike. Inhibition handles alerts that look different but share a cause: a node runs out of disk, and you also get evictions, crash loops and 5xx errors from everything scheduled on it. An inhibit rule mutes the target alerts while a source alert is firing, as long as the labels in equal match:

inhibit_rules:
  # A critical alert hides the warning version of the same alert.
  - source_matchers:
      - severity="critical"
    target_matchers:
      - severity="warning"
    equal: [alertname, cluster, service]

  # Disk pressure on a node hides pod symptoms on that same node.
  # Alert names are examples: use the names your rules actually produce.
  - source_matchers:
      - alertname="NodeDiskPressure"
    target_matchers:
      - alertname=~"KubePodEvicted|KubePodCrashLooping"
    equal: [cluster, node]

The second rule is where most inhibition goes wrong, so check two things.

Both alerts must carry every label in equal. Many pod-level alerts don't have a node label, because the metric they're built on doesn't. If the source has node="ip-10-2-14-7" and the target has no node at all, the values differ and nothing is inhibited. The fix is to add the label in the rule's expression (for example by joining with kube_pod_info), not to drop node from equal.

Missing labels count as equal. The Alertmanager docs say: "if all the label names listed in equal are missing from both the source and target alerts, the inhibition rule will apply." An equal list with a typo in a label name can therefore mute every matching target in every cluster. Keep the source and target matchers narrow, and test the rule (section 9).

Inhibited alerts aren't deleted. They still show in Prometheus and in the Alertmanager API as suppressed, with the inhibiting alert listed. That matters when you check you lost nothing.

5. Route non-urgent alerts to tickets, not phones

The cheapest change on this list: make paging opt-in. With the routing tree above, only severity="page" reaches the on-call receiver. Everything else lands in a ticket queue through a webhook:

receivers:
  - name: oncall-pager
    pagerduty_configs:
      - routing_key_file: /etc/alertmanager/secrets/pagerduty-key
  - name: ticket-queue
    webhook_configs:
      - url: https://alert-to-ticket.internal.example/hook
        send_resolved: true
  - name: team-channel
    slack_configs:
      - api_url_file: /etc/alertmanager/secrets/slack-webhook
        channel: '#ops-alerts'

The webhook service should create one ticket per alert group (use the payload's groupKey), comment on it on repeats instead of opening a new one, and close it when the alert resolves. That turns the cert-expiring-in-14-days alert from six pages a week into one ticket.

Check routing changes without deploying them: amtool check-config alertmanager.yml validates the file, and amtool config routes test --config.file=alertmanager.yml severity=page service=payments prints which receiver a given label set would reach.

6. Tune thresholds with for: and keep_firing_for:

A rule that fires for 40 seconds and resolves on its own is flapping. Two fields fix most of it:

groups:
  - name: payments
    rules:
      - alert: PaymentsHighErrorRate
        expr: |
          sum(rate(http_requests_total{job="payments", code=~"5.."}[5m]))
            / sum(rate(http_requests_total{job="payments"}[5m])) > 0.02
        for: 10m              # must be true for 10 minutes before it fires
        keep_firing_for: 5m   # stays firing 5 minutes after it stops being true
        labels:
          severity: page
        annotations:
          summary: Payments 5xx above 2% for 10 minutes
          runbook_url: https://runbooks.internal.example/payments-5xx

for keeps the alert in the pending state until the condition has held for that long, so short spikes never notify. keep_firing_for (Prometheus 2.42 and later) keeps it firing for a while after the condition clears, so a metric that dips below the threshold for one evaluation doesn't produce a resolve and a fresh page. Grafana alert rules have the same idea as a pending period.

Don't guess the durations. Use the firing-time query from section 1 and your incident history: if real incidents for this rule always lasted longer than 15 minutes and the noise always cleared in under 3, a for of 5 to 10 minutes costs you little detection time. Write the decision down as a unit test so the next person can see why:

# payments_test.yml, run with: promtool test rules payments_test.yml
rule_files:
  - payments_rules.yml
evaluation_interval: 1m
tests:
  - interval: 1m
    input_series:
      # 3-minute spike to 5% errors, then normal: must NOT fire.
      - series: 'http_requests_total{job="payments", code="500"}'
        values: '0+0x10 0+5x3 15+0x20'
      - series: 'http_requests_total{job="payments", code="200"}'
        values: '0+100x40'
    alert_rule_test:
      - eval_time: 20m
        alertname: PaymentsHighErrorRate
        exp_alerts: []

Add a second test with a sustained error rate that must fire. The pair stops anyone loosening the rule until it misses a real outage.

7. Which runbooks are safe to automate

If the response to an alert is always the same (clear old logs, restart one stuck worker, rotate one node), a person doing it at 2 a.m. adds delay and little judgement. But automatic remediation is only safe when three things are true. We call them verifiable, reversible and contained:

Wire it so the runbook runs, waits, reads the triggering metric again, and if the metric hasn't recovered, stops and pages a person with what it tried. An automated fix that fails quietly is worse than a page. Log every run where the team will see it, so the weekly review can count automated fixes too: a rule fixed automatically 30 times a week is a root cause waiting for a ticket.

8. The weekly noise review

Alert rules decay as the system under them changes, so tuning once doesn't hold. A 30-minute weekly slot with the people who were on call does. A format that works:

  1. The number (5 minutes). Alerts that reached a human this week versus last, and how many were acted on.
  2. The top five rules by pages (20 minutes). For each: how many pages, how many acted on, how long it fired. Decide one of: keep, retune, demote to ticket, automate, delete. Write down why.
  3. Owners (5 minutes). Every decision becomes a pull request against the rules repo with a named owner, merged by the following week.

Two rules keep it honest. Changes go through pull requests, never through a silence that someone forgets to remove. And nobody mutes a rule because it's annoying without saying what would catch the problem instead.

9. Check you didn't lose signal

You can always reduce alerts by alerting on less. The number from section 1 can't tell you whether you did that, so every change needs a second check.

10. Where products and agents fit

Everything above works with open-source Prometheus and Alertmanager, and for many teams it's enough. The commercial tools automate parts of it:

What none of them do for you is the judgement in sections 2, 7, 8 and 9: deciding which rule is wrong, fixing it in your repo, and proving nothing was lost. That's the part an AI agent built for your environment can take on. Lookout is a mock demo of that design. It groups alerts by shared labels, timing and cause, runs only pre-approved runbooks with limits and checks each fix against the triggering metric, attaches alerts to existing tickets, and runs this guide's weekly review itself, proposing rule changes as pull requests and never silencing a rule on its own. It's built around one number, alerts that reach a human. See this design working, or read how we approach custom autonomous systems.

11. FAQ

What is alert fatigue?

Alert fatigue is what happens when on-call engineers get so many alerts that lead nowhere that they stop treating each one as urgent. Pages get acknowledged without being read, and the alert that matters is missed or answered late. The fix is fewer alerts reaching people, not faster responses to all of them.

How do I reduce the number of alerts from Alertmanager?

Start by counting pages per alert rule for a few weeks and noting which ones anyone acted on. Then group alerts that describe one problem with group_by, inhibit symptom alerts while their cause is firing with inhibit_rules, route anything non-urgent to a ticket queue instead of the pager, and add for: durations to rules that flap. Review the noisiest rules weekly and change them by pull request.

What is the difference between grouping and inhibition in Alertmanager?

Grouping merges alerts that have the same values for the labels in group_by into one notification, for example nine crash-looping pods in one namespace. Inhibition mutes alerts of a different kind while a related source alert is firing, for example pod evictions on a node that is out of disk. Grouping reduces repeats; inhibition removes symptoms of a cause you already know about.

Why is my Alertmanager inhibit rule not working?

The usual cause is the equal list. Both the source and target alerts must have the same value for every label in it, and if the target alert doesn't carry a label such as node at all, the values differ and nothing is inhibited. A group_wait that's too short is the other cause: the target's notification goes out before the source alert arrives. Check the labels on both alerts in the Alertmanager UI or with amtool.

How long should the for: duration be on a Prometheus alert?

Long enough that short spikes which clear on their own never fire, and short enough that a real incident still pages in time. Look at how long this rule's noisy firings lasted compared with real incidents, and pick a value between them. Add keep_firing_for (Prometheus 2.42 and later) if the alert resolves and re-fires when the metric hovers around the threshold.

Which runbooks are safe to automate?

Runbooks whose result you can check with a metric, that you can undo, and that act on a small, bounded scope, such as clearing old logs on one node or restarting one stateless worker. The automation should re-read the metric that triggered the alert after it runs and page a person if the metric hasn't recovered. Anything that deletes data or fails over a database should stay with a human.

How do I know I didn't lose important alerts after tuning?

Replay recent real incidents against the changed rules, ideally as promtool unit tests, and confirm each one would still have paged. Compare what fired in Prometheus with what Alertmanager delivered for a week to catch inhibit rules that hide too much. Track incidents first reported by customers or other teams: if that number rises after a tuning round, you cut too deep.

Written by Ilia Dubovskii, founder of Beamreach. Config was checked against the Prometheus and Alertmanager documentation, and product details against each vendor's own site, on 10 October 2026. If something here no longer matches the docs, the docs win. Tell us and we'll fix this page.

Related: Lookout: AI alert triage demo · Autonomous DevOps · Custom autonomous systems · All guides for cloud teams