DR strategy spectrum
AWS defines four common disaster recovery (DR) patterns — exam questions often describe a scenario and expect you to name or choose the right tier.
| Strategy | RTO / RPO (typical) | Cost | Description |
|---|---|---|---|
| Backup and restore | Hours / hours | Lowest | Backups in S3; rebuild infra in DR Region on failure |
| Pilot light | Tens of minutes / minutes | Low–medium | Minimal core (DB replica, AMIs) running; scale up on disaster |
| Warm standby | Minutes / minutes | Medium | Scaled-down full stack always running in DR Region |
| Multi-site active/active | Near zero / near zero | Highest | Full production in multiple Regions serving traffic |
Map business RTO/RPO to strategy — there is no single "best" answer without requirements.
AWS Backup
Centralized backup for EC2, EBS, RDS, DynamoDB, EFS, FSx, Storage Gateway, etc.:
- Backup plans — schedule, retention, lifecycle to cold storage.
- Backup vault — encryption with KMS; access policies.
- Cross-Region copy — meet compliance and DR for backups.
- Legal hold and audit via CloudTrail.
Replaces one-off scripts when org needs consistent policy across services.
Amazon S3 for durability and DR
- 11 nines durability; Versioning + Cross-Region Replication (CRR) for backup buckets.
- S3 Object Lock — WORM for ransomware-resistant backups.
- Glacier tiers — long-term retention cheaper.
Store CloudFormation/Terraform templates + AMIs + config in S3 so backup and restore is repeatable.
AWS Elastic Disaster Recovery (DRS)
Continuous block-level replication of on-prem or cloud servers to staging area in AWS; drill failover to EC2 with minimal RPO. Exam alternative to DIY snapshot replication for lift-and-shift DR.
Route 53 in DR
- Failover routing — health-checked primary in Region A, secondary in Region B.
- Weighted shift — gradually move traffic during migration or DR test.
Combine with runbooks and regular game days.
Infrastructure as code
- Rebuild speed depends on automated provisioning (CloudFormation, CDK, Terraform).
- Pilot light often keeps RDS read replica or Aurora Global secondary + S3 static assets; run
terraform applyto scale compute on disaster.
Testing DR
- Game days — simulate Region loss.
- Backup restore drills — validate RPO with PITR restore to new instance.
- Monitor backup job failures in AWS Backup dashboard.
Exam traps
- Cheapest DR is backup and restore — wrong if question demands RTO in minutes.
- Active/active multi-Region solves global latency and DR but needs data conflict strategy (DynamoDB global tables, Aurora Global, etc.).
- Snapshots alone without automated redeploy still mean long RTO.
Official reference
SAA-C03 exam guide — disaster recovery and backup.