All challenges

Survive an AZ outage

advanced

Scenario & brief

This service runs two instances across two AZs with a Multi-AZ database — but if an entire AZ goes dark, the survivors can't absorb full peak, and the database has no read replica to fall back on.

Targets: 99.99% availability, critical durability, p99 ≤ 160 ms, under $1,300/mo. It should genuinely survive losing a whole AZ, not just one instance.

5,000 rps peakp99 ≤ 160ms99.99% availdurability: criticalbudget $1,300/mo

App Servers

Compute

2instances

System health

Erupting · SLA breach

28

/ 100

Score

SLA not met yet

Monthly cost

$653

Budget $1,300/mo · within budget

Metrics

Capacity34
Availability90
Durability25
Cost efficiency80

Requirements

  • Peak capacity 3600 rps compute · 1700 rps db (need ≥ 5000 rps)
  • p99 latency ~194 ms (need ≤ 160 ms)
  • Availability 99.95% (need ≥ 99.99%)
  • Durability at risk (need redundancy + backups)
  • Budget $653/mo (need ≤ $1300/mo)

Advisor

  • Compute tops out at ~3600 rps but peak demand is 5000 rps — requests will queue.
  • The database is saturated at peak — add a cache to shed read load, or scale it up.
  • Enable automated backups to satisfy the durability requirement.

Discussion

Sign in to join the discussion.

No comments yet. Be the first to start the discussion.

For learning purposes only. Costs and capacities are illustrative, not live AWS prices. Not affiliated with or endorsed by Amazon Web Services.