DevOps Incident Management

An incident is an unplanned event that harms a service, such as a website outage, a slow checkout page, or a failed payment system. Incident management is the set of habits that helps a team find the problem, fix it quickly, and learn from it.

Why a Process Helps

Think of a fire brigade. Firefighters do not argue about who holds the hose when the alarm rings. Each person knows a role, and the team follows a practiced routine. Software teams need the same calm structure. Without it, ten engineers shout in a chat room while the outage grows.

The Incident Lifecycle

 Detect --> Triage --> Respond --> Resolve --> Review
   |          |          |           |           |
 Alert     Rate the    Fix or     Service     Write a
 fires     impact      work       returns     postmortem
                       around     to normal

Severity Levels

Severity levels tell everyone how serious the incident is. Each company defines its own scale. A common version looks like this:

LevelMeaningExampleResponse
SEV1Major outageWebsite is down for all usersPage everyone, work until fixed
SEV2Serious damageCheckout fails for some usersPage on-call engineer
SEV3Minor damageOne report loads slowlyFix during work hours

Key Roles

Incident Commander

The commander leads the response. The commander assigns tasks, keeps the team focused, and makes final decisions. The commander does not need to fix the bug personally.

Technical Lead

The technical lead investigates the cause and guides the repair work.

Communications Lead

The communications lead updates the status page, support staff, and company leaders. Clear updates reduce panic and repeated questions.

Scribe

The scribe records a timeline of events, decisions, and commands during the incident. This record feeds the later review.

On-Call Duty

On-call engineers carry a phone or pager during a set shift. An alert wakes them when a serious problem appears. Healthy on-call programs rotate duty fairly, give rest after busy nights, and keep alerts meaningful. Too many false alarms exhaust people and hide real problems.

Runbooks

A runbook is a step-by-step guide for a known problem. A good runbook lists the symptoms, the checks to run, and the fix commands. Example runbook outline:

Problem: Web server disk is full
Check:   df -h
Action:  Delete old logs in /var/log/app
Verify:  Disk usage falls below 70 percent
Escalate: Call the platform team if usage stays high

Communication During an Incident

  • Open one shared channel for all incident talk.
  • Post updates at fixed times, such as every 30 minutes.
  • State facts, current impact, and the next update time.
  • Avoid guesses about the cause until the team confirms it.

Sample Timeline

 10:02  Alert fires: error rate above 20 percent
 10:05  On-call engineer acknowledges the page
 10:10  Team declares SEV2 and names a commander
 10:25  Recent release found as the likely cause
 10:30  Release rolled back
 10:38  Error rate returns to normal
 10:45  Incident closed, review scheduled

The team restored service by rolling back first and investigating later. Fast recovery matters more than a perfect diagnosis during the incident.

Blameless Postmortems

A postmortem is a written review after the incident. A blameless postmortem looks at systems and processes instead of pointing at a person. People share facts freely when they feel safe, and the team learns more. A postmortem usually contains:

  • A short summary and the user impact.
  • A timeline of events.
  • The root causes and contributing factors.
  • What worked well and what did not.
  • Action items with owners and due dates.

Escalation Policies

An escalation policy states who receives the alert and who receives it when the first person stays silent. A policy prevents an alert from dying in an unread inbox.

 Alert --> Primary on-call --(5 min, no reply)--> Secondary on-call
                                                       |
                                          (10 min, no reply)
                                                       v
                                               Engineering manager

Incident Metrics

Numbers show whether the team improves over time.

MetricFull NameWhat It Measures
MTTDMean Time to DetectTime from problem start to alert
MTTAMean Time to AcknowledgeTime from alert to human response
MTTRMean Time to RecoverTime from problem start to restored service

Use these numbers to find slow steps. A high MTTD points to weak monitoring. A high MTTR points to missing runbooks or hard rollbacks.

Status Pages

A status page tells customers what is happening. Short, honest updates build trust. Most pages use four stages: Investigating, Identified, Monitoring, and Resolved.

[Investigating] 10:10 - Some users cannot complete checkout. Our team is looking into it.
[Identified]    10:25 - A recent update caused the problem. We are reversing it.
[Monitoring]    10:38 - Checkout works again. We are watching the system.
[Resolved]      10:55 - The incident has ended.

Incident Tooling

  • Paging tools such as PagerDuty and Opsgenie alert the right person.
  • Chat tools such as Slack and Microsoft Teams host the incident channel.
  • Status page tools publish updates for customers.
  • Ticket tools track follow-up tasks from the review.

Game Days

A game day is a practice incident. The team invents a failure, such as a broken database, and works through the full response. Drills reveal missing permissions, outdated runbooks, and unclear roles at a time when nothing real is at risk.

Key Points

  • Clear roles and severity levels remove confusion.
  • Runbooks speed up repairs for known problems.
  • Rollback often beats live debugging.
  • Blameless reviews turn outages into lasting improvements.

Leave a Comment

Your email address will not be published. Required fields are marked *