Esta página ainda não está traduzida — é apresentada em inglês.

← Back to blog

What "99.9% uptime" actually counts

2026-10-05

engineeringstatus-pages

"99.9% uptime" sounds like a measurement of time. In okokumo it isn't one. It is a count: of all the probe results a check recorded in a window, the share that passed. Everything else about the number follows from that, and so do the ways it can mislead you.

This post goes through the formula as the code implements it, with worked numbers, so you can read a badge or a status page the way you would read any other metric: knowing what's in the numerator, what's in the denominator, and what got rounded away.

The formula

Every probe run writes one row to check_results with a success flag. An HTTP check retries once on a network error inside the same run, and a run still writes exactly one row. To build history we group those rows by UTC day (time_bucket('1 day', time)) and count two things per day: the total, and the number that passed.

The uptime percentage is the sum of passing results across the window divided by the sum of all results, times 100:

uptime = Σ up_count / Σ total × 100

Two things about this are easy to miss.

It is weighted by result, not by day or by minute. A day on which the check ran 1,440 times weighs 1,440 times more than a day with one result. If you tighten an interval from 5 minutes to 30 seconds halfway through the window, the recent half dominates the number.

There is no time in it. A failing result counts as one failure whether the outage behind it lasted four seconds or the whole interval. The number of probes is what limits the resolution, not the clock.

The window

The window is the last N calendar days in UTC, today included, and today has only run since midnight UTC. N depends on where the number appears:

  • Badges (the embeddable widget): 30 days, shown with one decimal (99.9%).
  • Status pages: 90 days, shown with two decimals (99.95% uptime).

Raw probe results are kept for 90 days, which is why the status page doesn't go further back.

How a day is classified

Each UTC day in the window gets one of four states, and the rules contain no thresholds:

  • up: every result that day passed.
  • down: every result that day failed. Not a single pass.
  • partial: anything in between. One failure in 1,440 is enough.
  • no data: no results at all that day.

A red day therefore means the check never passed even once, from 00:00 to 24:00 UTC. A bad afternoon is amber. And because the buckets are UTC, an outage from 23:50 to 00:10 UTC marks two days partial. For a team in Paris that was one evening; in UTC it was two dates.

Worked example: one 3-minute outage

Take an HTTP check on a 60-second interval. The scheduler adds ±10% jitter, so that's roughly 1,440 results a day. The service goes down for three minutes, and about three probes land in that gap and fail.

The day is partial: 1,437 passes and 3 failures. You see one amber square in a row of green.

The badge (30 days, about 43,200 results) shows 3 / 43,200 = 0.007% failed, so 99.993%, which one decimal rounds to 100.0%.

The status page (90 days, about 129,600 results) shows 99.998%, which rounds to 100.00%.

So the outage is visible in the day bars but invisible in the headline percentage. That's correct arithmetic, and you should know it before you quote the number. For the same check, the badge only drops below 100.0% at 22 failed results in 30 days. It then reads 99.9% anywhere from 22 to 64 failed results. The same "99.9%" covers outages whose lengths differ by a factor of three.

What 99.9% means in minutes

For a check that probes once a minute, 99.9% allows:

  • about 86 seconds a day,
  • 43.2 minutes in a 30-day month (44.6 in a 31-day one),
  • about 8.8 hours a year.

Those numbers only make sense if the interval is fine enough to see that much downtime. The free plan's minimum interval is 5 minutes, which gives 288 results a day and 8,640 in 30 days. A single failed probe then costs 0.012% on the badge, and an outage shorter than five minutes can fall between two probes and never be recorded. The interval sets the resolution. A 99.99% claim measured every 5 minutes is mostly a claim about where the probes happened to land.

Paused checks show no data

The scheduler only picks up enabled checks. A paused check stops producing results, so its days are no data. Those days add nothing to the numerator or the denominator, so a paused week neither drags the percentage down nor pads it.

That's deliberate. Pausing is not an outage, and a red bar would be a false statement to anyone reading your status page. But it cuts both ways: if you pause a check during an outage, the outage is never counted. That's why a maintenance window in okokumo keeps probing and only stops the paging (see Pause the alerts, not the check). Failures during maintenance do count against the percentage. The status page shows the component as "under maintenance" while the window is open, but nothing is subtracted from the history afterwards.

On a badge, a paused check shows "Monitoring paused", a grey ring and no percentage. On a status page the component's state reads as unknown, and the percentage still covers whatever days in the window have data.

No results: no number

If a check has no results anywhere in the window (a brand-new check, or one paused for the whole window), the percentage is hidden, not shown as zero. The badge shows – and the status page says "No data yet". Since 0 out of 0 has no meaningful value, 0% would mean "always down" and 100% would be a claim nothing backs.

The same reasoning applies to a young check. Three days of results on a 90-day status page means the percentage covers three days, and the other 87 bars are grey. Count the grey bars before you trust the decimals.

Heartbeats: what gets counted

For cron and heartbeat checks, results are written when a ping arrives: a success ping is a pass and a /fail ping is a failure. When a job goes silent, the check goes down and alerts you (The cron job that stopped running), but silence writes no result. A job that stopped running entirely shows grey no-data days, not red ones, and its percentage only reflects the runs that reported in. For a heartbeat, read the percentage as "share of reported runs that succeeded", and look at the check state to see whether it's running at all.

Uptime is stricter than alerting

Alerting waits for consecutive failures and, with more than one probe region active, confirmation from at least two regions (why). The percentage waits for nothing. Every failed result counts, including a one-region blip that never paged anyone. A day can be partial without any incident in your history. That's the number being literal, not contradictory.

Reading one as an engineer

When you see an okokumo percentage, on your own page or someone else's:

  1. Find the window. A badge covers 30 UTC days, a status page 90, and both include today so far.
  2. Look at the bars, not the decimals. Short outages round away in the percentage but not in the day colours. One amber square tells you more than "100.00%".
  3. Ask for the interval. A result is only as fine as the probe spacing.
  4. Count the grey. No-data days are excluded, not counted as up.
  5. Don't read the colour as the percentage. On a badge the ring colour is the check's current state. On a status page "degraded" means up right now, with at least one failure today.

The percentage is a summary for people who can't see your system. Your own incident history is still the record of what happened. For what to tell those people when the bars go amber, see Your status page is not a dashboard. More on how okokumo status pages work: Tower.