Design for failure
The SAA-C03 exam assumes everything fails eventually — AZ outages, instance corruption, bad deploys. Your architecture must detect, replace, and reroute automatically with minimal human intervention.
Availability zones and Regions
- Region — geographic area; fully independent failure domain from other Regions.
- Availability Zone (AZ) — one or more discrete data centers with redundant power and networking in a Region.
- Multi-AZ — run redundant components in at least two AZs in the same Region for HA within that Region.
- Multi-Region — disaster recovery or global low latency; higher complexity and cost.
Exam default: when a question says "highly available" without global users, think Multi-AZ in one Region first.
Elastic Load Balancing (ELB)
Distributes traffic across healthy targets:
| Type | Layer | Use case |
|---|---|---|
| ALB | Layer 7 (HTTP/HTTPS) | Microservices, path/host routing, WebSockets |
| NLB | Layer 4 (TCP/UDP) | Extreme performance, static IP, millions RPS |
| GLB | Gateway | Third-party appliances (IDS) |
| CLB | Legacy L4/L7 | Avoid for new designs |
Health checks — ELB only sends traffic to targets passing checks. Pair with Auto Scaling so unhealthy instances are replaced.
Cross-zone load balancing — ALB/NLB can distribute evenly across AZs (watch data transfer costs on NLB).
Amazon Route 53 routing policies
- Simple — one record, multiple values (random client choice).
- Weighted — split traffic for blue/green or canary (e.g. 90/10).
- Latency — lowest latency Region/AZ endpoint for user.
- Failover — primary + secondary health-checked records (active/passive DR).
- Geolocation / Geoproximity — route by user location or bias traffic to a Region.
- Multivalue answer — multiple healthy records; client tries another on failure.
Health checks can target endpoints, CloudWatch alarms, or calculated health of child checks.
EC2 Auto Scaling
- Scaling policies — target tracking (CPU, ALB request count), step scaling, scheduled.
- Launch template — AMI, instance type, user data, IAM role, multiple AZ subnets.
- Health check grace period — allow boot time before marking unhealthy.
- ReplaceUnhealthy — ASG terminates failed instances and launches new ones.
Pattern: ALB → ASG spanning minimum 2 AZs, desired capacity ≥ 2 for stateless web tier.
RTO and RPO (exam vocabulary)
- RPO (Recovery Point Objective) — max acceptable data loss measured in time (how far back you restore).
- RTO (Recovery Time Objective) — max acceptable downtime before service restored.
Lower RPO/RTO → more replication, more Regions, higher cost.
Stateless vs stateful tiers
- Stateless app tier — easy to scale horizontally behind ELB; session in ElastiCache/DynamoDB if needed.
- Stateful data tier — use managed Multi-AZ databases (RDS, Aurora) rather than self-managed failover on EC2.
Exam traps
- Single NAT Gateway is a single AZ failure point — use one NAT per AZ for production HA.
- Route 53 failover requires health checks on primary or it won't fail over.
- "Spread instances across AZs" alone is not enough without ELB + health checks + Auto Scaling.
Official reference
SAA-C03 exam guide — resilient architectures.