All challenges

Kill the single points of failure

beginner

Scenario & brief

This service runs a single app instance and a single database in one AZ, with no backups. It works — until anything fails.

Targets: 99.95% availability, critical durability, p99 ≤ 200 ms, under $1,200/mo. The latency budget is generous — this one is about resilience, not raw speed.

2,000 rps peakp99 ≤ 200ms99.95% availdurability: criticalbudget $1,200/mo

App Servers

Compute

1instances

System health

Erupting · SLA breach

28

/ 100

Score

SLA not met yet

Monthly cost

$653

Budget $1,200/mo · within budget

Metrics

Capacity100
Availability30
Durability25
Cost efficiency78

Requirements

  • Peak capacity 3600 rps compute · 3400 rps db (need ≥ 2000 rps)
  • p99 latency ~77 ms (need ≤ 200 ms)
  • Availability 99.00% (need ≥ 99.95%)
  • Durability at risk (need redundancy + backups)
  • Budget $653/mo (need ≤ $1200/mo)

Advisor

  • Compute runs in a single AZ — spread across ≥2 AZs (with ≥2 instances) to meet the availability target.
  • A single instance is a single point of failure — run at least two.
  • The database has no Multi-AZ standby or replica — a failure risks data loss. Enable Multi-AZ and backups.

Discussion

Sign in to join the discussion.

No comments yet. Be the first to start the discussion.

For learning purposes only. Costs and capacities are illustrative, not live AWS prices. Not affiliated with or endorsed by Amazon Web Services.