DevOps Incident Management
An incident is an unplanned event that harms a service, such as a website outage, a slow checkout page, or a failed payment system. Incident management is the set of habits that helps a team find the problem, fix it quickly, and learn from it.
Why a Process Helps
Think of a fire brigade. Firefighters do not argue about who holds the hose when the alarm rings. Each person knows a role, and the team follows a practiced routine. Software teams need the same calm structure. Without it, ten engineers shout in a chat room while the outage grows.
The Incident Lifecycle
Detect --> Triage --> Respond --> Resolve --> Review
| | | | |
Alert Rate the Fix or Service Write a
fires impact work returns postmortem
around to normal
Severity Levels
Severity levels tell everyone how serious the incident is. Each company defines its own scale. A common version looks like this:
| Level | Meaning | Example | Response |
|---|---|---|---|
| SEV1 | Major outage | Website is down for all users | Page everyone, work until fixed |
| SEV2 | Serious damage | Checkout fails for some users | Page on-call engineer |
| SEV3 | Minor damage | One report loads slowly | Fix during work hours |
Key Roles
Incident Commander
The commander leads the response. The commander assigns tasks, keeps the team focused, and makes final decisions. The commander does not need to fix the bug personally.
Technical Lead
The technical lead investigates the cause and guides the repair work.
Communications Lead
The communications lead updates the status page, support staff, and company leaders. Clear updates reduce panic and repeated questions.
Scribe
The scribe records a timeline of events, decisions, and commands during the incident. This record feeds the later review.
On-Call Duty
On-call engineers carry a phone or pager during a set shift. An alert wakes them when a serious problem appears. Healthy on-call programs rotate duty fairly, give rest after busy nights, and keep alerts meaningful. Too many false alarms exhaust people and hide real problems.
Runbooks
A runbook is a step-by-step guide for a known problem. A good runbook lists the symptoms, the checks to run, and the fix commands. Example runbook outline:
Problem: Web server disk is full
Check: df -h
Action: Delete old logs in /var/log/app
Verify: Disk usage falls below 70 percent
Escalate: Call the platform team if usage stays highCommunication During an Incident
- Open one shared channel for all incident talk.
- Post updates at fixed times, such as every 30 minutes.
- State facts, current impact, and the next update time.
- Avoid guesses about the cause until the team confirms it.
Sample Timeline
10:02 Alert fires: error rate above 20 percent 10:05 On-call engineer acknowledges the page 10:10 Team declares SEV2 and names a commander 10:25 Recent release found as the likely cause 10:30 Release rolled back 10:38 Error rate returns to normal 10:45 Incident closed, review scheduled
The team restored service by rolling back first and investigating later. Fast recovery matters more than a perfect diagnosis during the incident.
Blameless Postmortems
A postmortem is a written review after the incident. A blameless postmortem looks at systems and processes instead of pointing at a person. People share facts freely when they feel safe, and the team learns more. A postmortem usually contains:
- A short summary and the user impact.
- A timeline of events.
- The root causes and contributing factors.
- What worked well and what did not.
- Action items with owners and due dates.
Escalation Policies
An escalation policy states who receives the alert and who receives it when the first person stays silent. A policy prevents an alert from dying in an unread inbox.
Alert --> Primary on-call --(5 min, no reply)--> Secondary on-call
|
(10 min, no reply)
v
Engineering manager
Incident Metrics
Numbers show whether the team improves over time.
| Metric | Full Name | What It Measures |
|---|---|---|
| MTTD | Mean Time to Detect | Time from problem start to alert |
| MTTA | Mean Time to Acknowledge | Time from alert to human response |
| MTTR | Mean Time to Recover | Time from problem start to restored service |
Use these numbers to find slow steps. A high MTTD points to weak monitoring. A high MTTR points to missing runbooks or hard rollbacks.
Status Pages
A status page tells customers what is happening. Short, honest updates build trust. Most pages use four stages: Investigating, Identified, Monitoring, and Resolved.
[Investigating] 10:10 - Some users cannot complete checkout. Our team is looking into it.
[Identified] 10:25 - A recent update caused the problem. We are reversing it.
[Monitoring] 10:38 - Checkout works again. We are watching the system.
[Resolved] 10:55 - The incident has ended.Incident Tooling
- Paging tools such as PagerDuty and Opsgenie alert the right person.
- Chat tools such as Slack and Microsoft Teams host the incident channel.
- Status page tools publish updates for customers.
- Ticket tools track follow-up tasks from the review.
Game Days
A game day is a practice incident. The team invents a failure, such as a broken database, and works through the full response. Drills reveal missing permissions, outdated runbooks, and unclear roles at a time when nothing real is at risk.
Key Points
- Clear roles and severity levels remove confusion.
- Runbooks speed up repairs for known problems.
- Rollback often beats live debugging.
- Blameless reviews turn outages into lasting improvements.
