Uni Cert / SAA-C03 / High availability and fault tolerance

High availability and fault tolerance

Multi-AZ, load balancing, Route 53 routing policies, and recovery objectives.

Estimated reading: ~30 min

Design for failure

The SAA-C03 exam assumes everything fails eventually — AZ outages, instance corruption, bad deploys. Your architecture must detect, replace, and reroute automatically with minimal human intervention.

Availability zones and Regions

  • Region — geographic area; fully independent failure domain from other Regions.
  • Availability Zone (AZ) — one or more discrete data centers with redundant power and networking in a Region.
  • Multi-AZ — run redundant components in at least two AZs in the same Region for HA within that Region.
  • Multi-Region — disaster recovery or global low latency; higher complexity and cost.

Exam default: when a question says "highly available" without global users, think Multi-AZ in one Region first.

Elastic Load Balancing (ELB)

Distributes traffic across healthy targets:

TypeLayerUse case
ALBLayer 7 (HTTP/HTTPS)Microservices, path/host routing, WebSockets
NLBLayer 4 (TCP/UDP)Extreme performance, static IP, millions RPS
GLBGatewayThird-party appliances (IDS)
CLBLegacy L4/L7Avoid for new designs

Health checks — ELB only sends traffic to targets passing checks. Pair with Auto Scaling so unhealthy instances are replaced.

Cross-zone load balancing — ALB/NLB can distribute evenly across AZs (watch data transfer costs on NLB).

Amazon Route 53 routing policies

  • Simple — one record, multiple values (random client choice).
  • Weighted — split traffic for blue/green or canary (e.g. 90/10).
  • Latency — lowest latency Region/AZ endpoint for user.
  • Failover — primary + secondary health-checked records (active/passive DR).
  • Geolocation / Geoproximity — route by user location or bias traffic to a Region.
  • Multivalue answer — multiple healthy records; client tries another on failure.

Health checks can target endpoints, CloudWatch alarms, or calculated health of child checks.

EC2 Auto Scaling

  • Scaling policies — target tracking (CPU, ALB request count), step scaling, scheduled.
  • Launch template — AMI, instance type, user data, IAM role, multiple AZ subnets.
  • Health check grace period — allow boot time before marking unhealthy.
  • ReplaceUnhealthy — ASG terminates failed instances and launches new ones.

Pattern: ALB → ASG spanning minimum 2 AZs, desired capacity ≥ 2 for stateless web tier.

RTO and RPO (exam vocabulary)

  • RPO (Recovery Point Objective) — max acceptable data loss measured in time (how far back you restore).
  • RTO (Recovery Time Objective) — max acceptable downtime before service restored.

Lower RPO/RTO → more replication, more Regions, higher cost.

Stateless vs stateful tiers

  • Stateless app tier — easy to scale horizontally behind ELB; session in ElastiCache/DynamoDB if needed.
  • Stateful data tier — use managed Multi-AZ databases (RDS, Aurora) rather than self-managed failover on EC2.

Exam traps

  • Single NAT Gateway is a single AZ failure point — use one NAT per AZ for production HA.
  • Route 53 failover requires health checks on primary or it won't fail over.
  • "Spread instances across AZs" alone is not enough without ELB + health checks + Auto Scaling.

Official reference

SAA-C03 exam guide — resilient architectures.

Official reference: AWS documentation

Chat with G.U.S.

Share suggestions to improve the site or any complaints. We use your feedback to make unigrat.com better.

Hi, I am G.U.S. — Growth Upgrade Suggestions.

Share your suggestions or complaints below.