Availability budget — regional API against a 99.99 % monthly SLO
Reliability block diagram of a three-AZ regional payments API stated in MTBF and MTTR, computing composite availability against a 99.99 % monthly SLO and ranking which block owns the failure — the arithmetic an SRE or architect must show before committing to an availability target.
Make it your own.
title "Payments API, eu-west-1 — availability budget for a 99.99 % monthly SLO"
mission 720h
# Mission is one 30-day SLA window, because that is the window the
# commitment is written in. Each block is stated as MTBF/MTTR from the
# incident record, so the availability under the figure is auditable:
# MTTR is time to RESTORE service, not time to notice.
block dns "Route 53 alias record" mtbf: 1000000h mttr: 0.50h
block alb "Application Load Balancer" mtbf: 17520h mttr: 0.25h
# One serving stack per zone: 6 Fargate tasks behind that zone's NAT
# gateway. 1/8760 h + 1/43800 h combines to a stack MTBF of 7300 h.
block az-a "Serving stack — eu-west-1a" mtbf: 7300h mttr: 0.10h
block az-b "Serving stack — eu-west-1b" mtbf: 7300h mttr: 0.10h
block az-c "Serving stack — eu-west-1c" mtbf: 7300h mttr: 0.10h
# Aurora: writer in 1a, reader in 1b promoted on failure. The switch
# probability is the observed success rate of automatic failover.
block aur-w "Aurora writer — 1a" mtbf: 4380h mttr: 0.25h
block aur-r "Aurora reader — 1b (promoted)" mtbf: 4380h mttr: 0.25h
# Regional dependencies with no redundancy this team owns.
block kms "KMS customer-managed key" mtbf: 87600h mttr: 0.50h
block sm "Secrets Manager" mtbf: 87600h mttr: 0.50h
series {
dns
alb
kofn 2 of 3 { az-a; az-b; az-c }
standby hot { primary: aur-w; spare: aur-r; switch: 0.995 }
kms
sm
}