DevOps Backup and Disaster Recovery

A backup is a saved copy of data. Disaster recovery is the plan that brings systems back after a major failure, such as a data center fire, a cloud region outage, a ransomware attack, or a mistaken deletion. Backups protect the data, and disaster recovery protects the whole service.

Why Both Matter

Think of important home documents. You keep photocopies in a drawer and spare keys with a trusted neighbor. A flood still strikes, so you know where to stay and who to call. Backups work like the photocopies. The recovery plan works like the written instructions for the bad day.

Two Numbers Every Team Sets

RPO: Recovery Point Objective

RPO states how much data loss the business accepts. An RPO of one hour means the team can lose at most the last hour of changes. A shorter RPO needs more frequent backups.

RTO: Recovery Time Objective

RTO states how long the service may stay down. An RTO of four hours means the team must restore service within four hours. A shorter RTO needs more ready-to-use equipment.

 Last good   Disaster                Service
 backup      strikes                restored
    |           |                      |
 ---+-----------+----------------------+------> time
    |<-- RPO -->|<------- RTO ------->|
    data loss        downtime

Types of Backups

TypeWhat It SavesSpeedRestore
FullEverythingSlowSimple and fast
IncrementalChanges since the last backup of any kindFastNeeds the full backup and every increment
DifferentialChanges since the last full backupMediumNeeds the full backup and the latest differential

The 3-2-1 Rule

  3 copies of your data
  2 different storage types (disk, cloud, tape)
  1 copy stored off-site

  [ Live data ] --> [ Local backup ] --> [ Cloud backup in another region ]

The rule protects against a single disk failure, a site failure, and human error. Add an immutable copy that nobody can change or delete for a set period. An immutable copy survives ransomware.

Disaster Recovery Strategies

StrategyDescriptionRecovery TimeCost
Backup and RestoreRebuild from stored backupsHours to daysLowest
Pilot LightKeep core pieces such as the database running in a second regionTens of minutes to hoursLow
Warm StandbyRun a smaller full copy in a second regionMinutesMedium
Multi-Site ActiveRun full copies in two regions at onceNear zeroHighest

Example: Database Backup Script

#!/bin/bash
date=$(date +%F)
pg_dump shopdb | gzip > /backups/shopdb-$date.sql.gz
aws s3 cp /backups/shopdb-$date.sql.gz s3://company-backups/db/
find /backups -name "*.sql.gz" -mtime +7 -delete

The script dumps the database, compresses the file, and copies it to cloud storage. The last line removes local copies older than seven days. A scheduler such as cron runs the script every night.

Infrastructure as Code Helps Recovery

Teams that describe servers and networks in Terraform or Ansible files rebuild an environment from code. Data still needs backups, and the code rebuilds everything around the data.

Test Your Restores

An untested backup is only a hope. Schedule regular restore drills in a separate environment. Measure the real recovery time and compare it with the RTO. Record the problems that appear and fix them before a real disaster arrives.

Encryption and Retention

Backups hold the same sensitive data as the live system, so encrypt them before storage and protect the keys separately. A retention policy decides how long each backup lives. Keeping every backup forever wastes money, and deleting backups too early removes recovery choices. Rules from law or contracts may set minimum periods.

Backup FrequencyKeep For
Daily7 days
Weekly4 weeks
Monthly12 months

Replication and Point-in-Time Recovery

Replication copies data to a second database almost instantly. A replica protects against hardware failure. A replica copies mistakes too, so a deleted table disappears from both copies. Point-in-time recovery solves this problem. The database saves a log of every change, and an engineer replays the log up to the moment just before the mistake.

 Full backup (Sunday) + change logs (Mon, Tue, Wed ... 14:31)
                                                   ^
                                  Restore to 14:30, before the bad delete

Failover and Failback

Failover moves traffic from the failed site to the recovery site. Failback returns traffic to the original site after repair. Both steps need a written plan and a practiced team.

 Normal:    Users --> Primary Region        Recovery Region (idle or small)
 Failover:  Users --> Recovery Region       Primary Region (down)
 Failback:  Users --> Primary Region        Recovery Region (idle again)

DNS changes, database promotion, and cache warm-up make up the usual failover steps. Automate as many steps as possible to cut the human error that appears under stress.

Kubernetes Backups

A cluster holds two kinds of state: object definitions stored in etcd and data stored on persistent volumes. The open source tool Velero backs up both.

velero backup create daily-backup --include-namespaces shop
velero restore create --from-backup daily-backup

Teams that keep manifests in Git rebuild the object definitions from the repository and use Velero or volume snapshots for the data.

Backup Monitoring

Backup jobs fail silently when nobody watches them. Send a success or failure signal to the monitoring system after every job. Alert on missing backups, shrinking backup sizes, and slow jobs. A backup that stopped three weeks ago is an unpleasant surprise during a disaster.

Key Points

  • RPO limits data loss, and RTO limits downtime.
  • The 3-2-1 rule gives a reliable backup pattern.
  • Higher recovery speed costs more money.
  • Regular restore tests prove that the plan works.

Leave a Comment

Your email address will not be published. Required fields are marked *