Esta página aún no está traducida — se muestra en inglés.

← Back to blog

Verifying an okokumo webhook — the signature, the retries, and the duplicate you'll eventually get

2026-09-28

guidealerting

A webhook endpoint is a public URL that takes a POST from anyone. If yours pages someone or opens a ticket, you need to know the request came from us, and you need to know what we do when your endpoint is slow or down, because that decides whether you see an alert zero times, once or twice. This post covers verifying the signature, how retries work, and making a duplicate harmless. Every claim here was checked against the code that sends the webhook.

What arrives

A webhook channel POSTs JSON to your HTTPS URL when a check goes down, when it recovers, and when you press test send. It also fires in two cases that are not state changes: a check paused by a plan change, and a check still down when its maintenance window ends. The request carries three headers:

Content-Type: application/json
User-Agent: SentinelAlerts/0.1
X-Sentinel-Signature: sha256=60d64e7348…

and a body like this one:

{
  "transition_id": "9c3f6c0a-1b2d-4e5f-8a9b-0c1d2e3f4a5b",
  "organization_id": "d649aac4-…",
  "check_id": "a59a12c1-…",
  "check_name": "nightly-backup",
  "from_state": "up",
  "to_state": "down",
  "reason": "the job reported a failure",
  "failure_body": "pg_dump: error: \u003cdb\u003e \u0026 café\n",
  "occurred_at": "2026-09-28T03:12:00Z",
  "link": "https://app.okokumo.com/checks/a59a12c1-…"
}

to_state: "down" is an incident and "up" is a recovery. test: true appears only on test sends, and paused: true only on pause notices. failure_body is present only when a heartbeat job reported its own failure. It is that job's raw stderr, so treat it as untrusted input.

Look closely at failure_body above. The job wrote <db> & café. The body carries \u003cdb\u003e \u0026 café, because the sender's JSON encoder escapes <, > and &. That detail is what trips up the most common verification mistake.

The signature is over bytes, not over JSON

The header value is sha256= followed by the lowercase hex HMAC-SHA256 of the exact request body. The key is your channel's signing secret, used as the string it is. You can set your own secret when you create the channel through the API (config.secret). If you leave it out, we generate 64 hex characters. Either way, the API returns the secret in the channel's config. The generated secret is also used as a string, so don't hex-decode it first. Editing the channel keeps the secret. Creating a new channel gets a new one.

The usual mistake: your framework parses the JSON, you call JSON.stringify(req.body) or json.dumps(payload) to get "the body" back, and you sign that. You get a document that means the same thing but has different bytes: < in place of \u003c, different spacing, maybe a different key order. The HMAC doesn't match, and you return 401 on every real alert. We ran the payload above through all three verifiers below. Each one accepts the raw body. Each one rejects the same body after a parse-and-re-serialise round trip.

So read the raw body before anything parses it, verify, and only then call JSON.parse.

Node (with Express, mount express.raw({ type: 'application/json' }) on this route in place of express.json(), so req.body is a Buffer):

import crypto from 'node:crypto';

export function verify(rawBody, header, secret) {
	const expected = 'sha256=' + crypto.createHmac('sha256', secret).update(rawBody).digest('hex');
	const a = Buffer.from(expected);
	const b = Buffer.from(header ?? '');
	return a.length === b.length && crypto.timingSafeEqual(a, b);
}

The length check is there because timingSafeEqual throws on buffers of different lengths instead of returning false. The length of a SHA-256 hex digest is public anyway, so checking it first leaks nothing.

Python (Flask: request.get_data(); FastAPI: await request.body()):

import hashlib
import hmac

def verify(raw_body: bytes, header: str | None, secret: str) -> bool:
    expected = "sha256=" + hmac.new(secret.encode(), raw_body, hashlib.sha256).hexdigest()
    return hmac.compare_digest(expected, header or "")

Go (io.ReadAll(r.Body), then verify, then json.Unmarshal):

func verify(rawBody []byte, header, secret string) bool {
	mac := hmac.New(sha256.New, []byte(secret))
	mac.Write(rawBody)
	expected := "sha256=" + hex.EncodeToString(mac.Sum(nil))
	return hmac.Equal([]byte(expected), []byte(header))
}

All three use a constant-time comparison. With a plain ==, the time the compare takes depends on how many leading characters match. Over enough requests, that difference is how an attacker could work out a valid signature byte by byte. timingSafeEqual, compare_digest and hmac.Equal exist to close that hole.

One limit to know: the signature covers the body only. There is no signed timestamp, so it proves who sent a request but not when. Someone who captures one delivery could send it again. The deduplication below handles that too.

Retries: what your status code tells us

Each delivery gets up to three attempts: one right away, one a second later, and a last one three seconds after that. Each attempt has a 10-second timeout. What happens next depends on what you return:

  • 2xx: delivered. We make no more attempts.
  • 3xx: the sender follows the redirect, and whatever comes back from the new URL decides the result. A 301, 302 or 303 turns the POST into a GET with no body, and if that GET returns 200 we count it as delivered even though your handler never saw the alert. A 307 or 308 keeps the POST and its signature. Don't put a redirect in front of the endpoint; configure the final URL.
  • 4xx, any of them: final. That includes 401, 404 and 429. We don't retry a rejection, because a bad signature or a wrong path won't fix itself in four seconds.
  • 5xx, a timeout, or a connection error: we try again, up to the third attempt.

This has two consequences. First, return 401 when verification fails. It's the honest answer and it stops the retries. But a verifier that rejects everything, like the re-serialising one above, drops every alert without a trace on your side. The delivery log on the check's page will show it as failed. The quickest way to catch it is the test send, since a webhook channel isn't confirmed until a test delivery succeeds. Second, a 429 from your rate limiter is final too. If your endpoint sits behind one, exempt this route.

The duplicate you'll eventually get

The retry loop re-sends the same bytes. Say your handler writes the alert to your database and then takes 11 seconds to respond, or your load balancer returns a 502 after the app has already committed. From where we sit, that attempt failed. We send it again, and you process it twice. Nothing is broken when that happens. It's what at-least-once delivery with retries means. So:

Acknowledge fast. Verify, store, return 200, and do the slow work somewhere else. Most duplicates come from a handler that is slower than the 10-second timeout.

Deduplicate on transition_id. A down or recovery alert carries the ID of the state change that caused it, so the same key covers all three attempts of that delivery. Insert it into a table with a unique constraint, and when the insert conflicts, return 200 and do nothing. A 200 on a duplicate is correct. A 409 would work too, since any 4xx stops retries, but it would show as a failure in your delivery log.

One exception. Pause notices and still-down-after-maintenance notices are not state changes, so they have no transition. For those, transition_id is the all-zeros UUID. If you dedupe on it naively, the second pause notice you ever get looks like a duplicate of the first. For those two cases, key on the signature header. We sign the body once per send and resend it byte for byte, so every retry carries the same signature, while two different notices differ at least in occurred_at. The key also blocks replays for as long as you keep it: a captured request sent again is just another duplicate.

Test sends get a fresh random transition_id each time, so they never collide with real alerts. Filter on test if you don't want them to reach anyone.

The short version

Read the raw body. Compute sha256= plus the hex HMAC of those bytes, keyed with the secret string, and compare in constant time. Return 401 on a mismatch and 200 quickly on a match. Store transition_id, or the signature when the ID is all zeros, under a unique constraint, and treat a conflict as success. Then press test send and confirm the channel goes from pending to confirmed.