Skip to content
Reliability block diagram templates

Availability budget — regional API against a 99.99 % monthly SLO

Reliability block diagram of a three-AZ regional payments API stated in MTBF and MTTR, computing composite availability against a 99.99 % monthly SLO and ranking which block owns the failure — the arithmetic an SRE or architect must show before committing to an availability target.

Template previewReliability block diagram
Payments API, eu-west-1 — availability budget for a 99.99 % monthly SLOMission 720 h · 30.0 d · 9 blocks · 2 redundant arrangements · R 0.898276Route 53 alias recordR 0.999280 · A 1.00000Application Load BalancerR 0.959737 · A 0.999992 of 3 · R 0.975193Serving stack — eu-west-1aR 0.906078 · A 0.99999Serving stack — eu-west-1bR 0.906078 · A 0.99999Serving stack — eu-west-1cR 0.906078 · A 0.99999hot standby · switch 0.995 · R 0.976379Aurora writer — 1aR 0.848417 · A 0.99994Aurora reader — 1b (promoted)R 0.848417 · A 0.99994KMS customer-managed keyR 0.991815 · A 0.99999Secrets ManagerR 0.991815 · A 0.99999System reliability0.898276Unreliability0.1017System MTTF3 142 h · 131 dAvailability0.999974Reliability importance — Birnbaum ∂R/∂Rᵢ, with criticality share1Application Load Balancer0.935960crit 37%2KMS customer-managed key0.905689crit 7.29%3Secrets Manager0.905689crit 7.29%4Route 53 alias record0.898923crit 0.64%5Serving stack — eu-west-1a0.156777crit 14%6Serving stack — eu-west-1b0.156777crit 14%7Serving stack — eu-west-1c0.156777crit 14%8Aurora writer — 1a0.143361crit 21%+1 more blocks not shownWeakest link: Application Load Balancer — R 0.959737, Birnbaum 0.935960, 37% of system failure. Improving this one block moves the system number further than improving any other.Redundancy check: removing any one of the 4 branches makes its own group measurably more likely to fail — the least of them by 542%. Every one of them is earning its place.

Make it your own.

title "Payments API, eu-west-1 — availability budget for a 99.99 % monthly SLO"
mission 720h

# Mission is one 30-day SLA window, because that is the window the
# commitment is written in. Each block is stated as MTBF/MTTR from the
# incident record, so the availability under the figure is auditable:
# MTTR is time to RESTORE service, not time to notice.

block dns   "Route 53 alias record"            mtbf: 1000000h mttr: 0.50h
block alb   "Application Load Balancer"        mtbf: 17520h   mttr: 0.25h

# One serving stack per zone: 6 Fargate tasks behind that zone's NAT
# gateway. 1/8760 h + 1/43800 h combines to a stack MTBF of 7300 h.
block az-a  "Serving stack — eu-west-1a"       mtbf: 7300h    mttr: 0.10h
block az-b  "Serving stack — eu-west-1b"       mtbf: 7300h    mttr: 0.10h
block az-c  "Serving stack — eu-west-1c"       mtbf: 7300h    mttr: 0.10h

# Aurora: writer in 1a, reader in 1b promoted on failure. The switch
# probability is the observed success rate of automatic failover.
block aur-w "Aurora writer — 1a"               mtbf: 4380h    mttr: 0.25h
block aur-r "Aurora reader — 1b (promoted)"    mtbf: 4380h    mttr: 0.25h

# Regional dependencies with no redundancy this team owns.
block kms   "KMS customer-managed key"         mtbf: 87600h   mttr: 0.50h
block sm    "Secrets Manager"                  mtbf: 87600h   mttr: 0.50h

series {
  dns
  alb
  kofn 2 of 3 { az-a; az-b; az-c }
  standby hot { primary: aur-w; spare: aur-r; switch: 0.995 }
  kms
  sm
}