DevOps Backup and Disaster Recovery
A backup is a saved copy of data. Disaster recovery is the plan that brings systems back after a major failure, such as a data center fire, a cloud region outage, a ransomware attack, or a mistaken deletion. Backups protect the data, and disaster recovery protects the whole service.
Why Both Matter
Think of important home documents. You keep photocopies in a drawer and spare keys with a trusted neighbor. A flood still strikes, so you know where to stay and who to call. Backups work like the photocopies. The recovery plan works like the written instructions for the bad day.
Two Numbers Every Team Sets
RPO: Recovery Point Objective
RPO states how much data loss the business accepts. An RPO of one hour means the team can lose at most the last hour of changes. A shorter RPO needs more frequent backups.
RTO: Recovery Time Objective
RTO states how long the service may stay down. An RTO of four hours means the team must restore service within four hours. A shorter RTO needs more ready-to-use equipment.
Last good Disaster Service
backup strikes restored
| | |
---+-----------+----------------------+------> time
|<-- RPO -->|<------- RTO ------->|
data loss downtime
Types of Backups
| Type | What It Saves | Speed | Restore |
|---|---|---|---|
| Full | Everything | Slow | Simple and fast |
| Incremental | Changes since the last backup of any kind | Fast | Needs the full backup and every increment |
| Differential | Changes since the last full backup | Medium | Needs the full backup and the latest differential |
The 3-2-1 Rule
3 copies of your data 2 different storage types (disk, cloud, tape) 1 copy stored off-site [ Live data ] --> [ Local backup ] --> [ Cloud backup in another region ]
The rule protects against a single disk failure, a site failure, and human error. Add an immutable copy that nobody can change or delete for a set period. An immutable copy survives ransomware.
Disaster Recovery Strategies
| Strategy | Description | Recovery Time | Cost |
|---|---|---|---|
| Backup and Restore | Rebuild from stored backups | Hours to days | Lowest |
| Pilot Light | Keep core pieces such as the database running in a second region | Tens of minutes to hours | Low |
| Warm Standby | Run a smaller full copy in a second region | Minutes | Medium |
| Multi-Site Active | Run full copies in two regions at once | Near zero | Highest |
Example: Database Backup Script
#!/bin/bash
date=$(date +%F)
pg_dump shopdb | gzip > /backups/shopdb-$date.sql.gz
aws s3 cp /backups/shopdb-$date.sql.gz s3://company-backups/db/
find /backups -name "*.sql.gz" -mtime +7 -deleteThe script dumps the database, compresses the file, and copies it to cloud storage. The last line removes local copies older than seven days. A scheduler such as cron runs the script every night.
Infrastructure as Code Helps Recovery
Teams that describe servers and networks in Terraform or Ansible files rebuild an environment from code. Data still needs backups, and the code rebuilds everything around the data.
Test Your Restores
An untested backup is only a hope. Schedule regular restore drills in a separate environment. Measure the real recovery time and compare it with the RTO. Record the problems that appear and fix them before a real disaster arrives.
Encryption and Retention
Backups hold the same sensitive data as the live system, so encrypt them before storage and protect the keys separately. A retention policy decides how long each backup lives. Keeping every backup forever wastes money, and deleting backups too early removes recovery choices. Rules from law or contracts may set minimum periods.
| Backup Frequency | Keep For |
|---|---|
| Daily | 7 days |
| Weekly | 4 weeks |
| Monthly | 12 months |
Replication and Point-in-Time Recovery
Replication copies data to a second database almost instantly. A replica protects against hardware failure. A replica copies mistakes too, so a deleted table disappears from both copies. Point-in-time recovery solves this problem. The database saves a log of every change, and an engineer replays the log up to the moment just before the mistake.
Full backup (Sunday) + change logs (Mon, Tue, Wed ... 14:31)
^
Restore to 14:30, before the bad delete
Failover and Failback
Failover moves traffic from the failed site to the recovery site. Failback returns traffic to the original site after repair. Both steps need a written plan and a practiced team.
Normal: Users --> Primary Region Recovery Region (idle or small) Failover: Users --> Recovery Region Primary Region (down) Failback: Users --> Primary Region Recovery Region (idle again)
DNS changes, database promotion, and cache warm-up make up the usual failover steps. Automate as many steps as possible to cut the human error that appears under stress.
Kubernetes Backups
A cluster holds two kinds of state: object definitions stored in etcd and data stored on persistent volumes. The open source tool Velero backs up both.
velero backup create daily-backup --include-namespaces shop
velero restore create --from-backup daily-backupTeams that keep manifests in Git rebuild the object definitions from the repository and use Velero or volume snapshots for the data.
Backup Monitoring
Backup jobs fail silently when nobody watches them. Send a success or failure signal to the monitoring system after every job. Alert on missing backups, shrinking backup sizes, and slow jobs. A backup that stopped three weeks ago is an unpleasant surprise during a disaster.
Key Points
- RPO limits data loss, and RTO limits downtime.
- The 3-2-1 rule gives a reliable backup pattern.
- Higher recovery speed costs more money.
- Regular restore tests prove that the plan works.
