DevOps Alerting
An alert is an automatic message that tells a person a system needs attention. Good alerting wakes the right engineer for a real problem and stays quiet otherwise. Poor alerting floods inboxes until nobody reads them.
The Smoke Detector Example
A good smoke detector rings when smoke fills a room and stays silent during normal cooking. A bad detector rings for every piece of toast. Households with a noisy detector remove its battery, and the next real fire goes unnoticed. Software teams face the same danger, called alert fatigue.
Monitoring, Alerting, and Dashboards
| Part | Job |
|---|---|
| Monitoring | Collects numbers and logs |
| Dashboards | Show the numbers to people who look |
| Alerting | Pushes a message when numbers go wrong |
The Alert Lifecycle
Metric crosses limit --> Rule waits (for: 5m) --> Alert fires
|
v
Alert resolves <-- Fix applied <-- Engineer acts <-- Notification sent
The waiting period filters short spikes. A CPU jump lasting ten seconds rarely needs a human, and a jump lasting twenty minutes does.
Symptoms Over Causes
Alert on what users feel, such as slow pages and failed checkouts. High CPU on one server may mean nothing to customers. Cause-based alerts, such as a full disk, still help when the disk will fill soon. Keep those alerts for problems that need action before users suffer.
Rules for a Good Alert
- Actionable: A person can do something about it right now.
- Urgent: The problem cannot wait until morning.
- Clear: The message states what broke, where, and how bad.
- Linked: The message holds a link to a dashboard and a runbook.
An alert that fails any rule belongs in a ticket or a daily report instead of a page.
Severity and Routing
| Severity | Example | Delivery |
|---|---|---|
| Critical | Checkout fails for all users | Phone call or page at any hour |
| Warning | Disk reaches 80 percent | Chat message in work hours |
| Info | Deployment finished | Log or ticket |
Static Thresholds and Smarter Methods
A static threshold fires at a fixed value, such as errors above 5 percent. Static limits are simple and easy to explain. Traffic patterns change by hour and by day, so some teams use anomaly detection, which learns normal behavior and flags large departures.
SLO Burn-Rate Alerts
A service level objective (SLO) sets a target, such as 99.9 percent of requests succeeding. The error budget is the allowed failure: 0.1 percent. A burn-rate alert fires when the service spends the budget too fast. A fast burn pages someone at once. A slow burn opens a ticket. This method links alerts to real customer impact.
Error budget: 100% Fast burn: [########----] gone in 2 hours --> page now Slow burn: [##----------] gone in 10 days --> open ticket
Alertmanager Routing
Alertmanager receives alerts from Prometheus, groups them, and sends them to the right place.
route:
receiver: team-chat
group_by: [alertname, service]
group_wait: 30s
repeat_interval: 4h
routes:
- match:
severity: critical
receiver: on-call-pager
receivers:
- name: team-chat
slack_configs:
- channel: "#ops-alerts"
- name: on-call-pager
pagerduty_configs:
- service_key: "<key>"Grouping
Grouping bundles many related alerts into one message. One failed database can trigger fifty alerts, and the group delivers them as a single page.
Inhibition
Inhibition mutes lower-level alerts when a higher-level alert already explains the problem. A "data center down" alert silences the "server unreachable" alerts inside it.
Silences
A silence mutes a known alert for a set time, such as during planned maintenance.
Reviewing and Testing Alerts
Hold a short alert review every month. Count how often each alert fired and how often it led to action. Delete or tune alerts that fire often without action. Test the full chain, from rule to phone call, with a fake alert before relying on it.
Key Points
- Every page should need a human and demand quick action.
- Symptom-based alerts match real user pain.
- Grouping, inhibition, and silences cut noise.
- Regular reviews keep alerting healthy.
