Concepts

What 99.9% Uptime Actually Means (And Why Your SLA Might Be Meaningless)

MontyJan 21, 2025 5 min read

Everyone quotes the nines. Fewer people can say what a number was measured against, and that second part is where all the meaning lives.

The table

AvailabilityDowntime per yearPer month
99%3d 15h7h 18m
99.9%8h 46m43m
99.99%52m4m 23s
99.999%5m 15s26s

Ninety-nine percent sounds respectable and allows nearly four days of downtime a year. And 99.99% allows four minutes and twenty-three seconds per month — less time than it takes most on-call engineers to open a laptop. Four nines requires that most failures are handled without a human in the loop at all.

The question the number does not answer

Availability is a ratio of good time to total time. Everything contentious hides in the definition of "good." A service where every request succeeds but takes eleven seconds is up by a naive check and unusable by any user's judgement. A service broken in one region reads 100% if your monitoring lives elsewhere. A broken checkout with a healthy homepage reports perfect availability through a total revenue outage.

The four questions

What was measured — a static health endpoint or a real user journey? From where — one location or many? What counted as down — the timeout threshold is a policy decision that silently determines your number. Over what window — annual figures smooth away incidents that were catastrophic for the people who lived through them.

SLI, SLO and SLA, briefly

An SLI is the measurement. An SLO is your internal target. An SLA is a contract with consequences, and it should always be looser than your SLO. A common structure is an SLA of 99.9% and an SLO of 99.95%, giving yourself twice the room internally that you promise externally.

Error budgets are the useful part

At 99.9% monthly you have about forty-four minutes to spend. Framed as a budget, it changes conversations: if you have used two minutes, you can afford a risky migration; if you have used forty, you should not ship anything discretionary. This is also the honest answer to "why can't we target 100%?" — a zero error budget means you can never change anything.

Measure the journey, not the endpoint

Point your availability measurement at something a user actually does. The number will be lower. It will also be true, and a true 99.7% is worth more than a fictional 99.99% every time an incident happens.

Monty