Engineering

Webhooks that don't lose events

Aug 25, 2026 · 2 min read

Every webhook system looks the same in the demo. An event happens, a POST arrives, everyone claps. The differences show up three weeks later, when a consumer was down for forty minutes and you find out whether those events were retried, dropped, or delivered twice into a handler that wasn't ready for that.

Delivery is a queue problem, not an HTTP problem

The naive version calls the consumer's endpoint inline, right where the event happens. It works until the consumer is slow, and then your API is slow too, because it's waiting on someone else's server. The fix is old and boring: write the event to a durable queue first, return, and let a dispatcher deliver it separately.

Horato runs deliveries through a dispatch queue that is drained every two minutes, with each delivery attempt logged. That log matters as much as the queue. When a customer asks why their handler didn't fire, you want to answer with a timestamped attempt and a response code, not a shrug.

Retries need a backoff and a limit

A consumer that returns 500 once probably deploys in a minute. A consumer that returns 500 for six hours is down, and hammering it every ten seconds helps nobody. Exponential backoff with a cap handles both cases with one policy.

  • Retry on 5xx and on timeouts. A timeout is a failure even though you never saw a status code.
  • Treat 4xx as permanent. If the consumer says the signature is wrong or the endpoint is gone, retrying won't change its mind.
  • Cap total attempts and surface dead deliveries in the dashboard so a human can replay them after the outage.

Sign everything, and make verification easy

An unsigned webhook endpoint is a public API that writes to your database. Anyone who finds the URL can post fake events into it. The standard answer is an HMAC signature over the raw body, sent in a header, verified with a shared secret.

The part most platforms get wrong is the consumer experience. If verifying a signature takes more than five lines of code, people skip it. Publish the exact algorithm, provide a helper in your SDKs, and include a timestamp in the signed payload so replayed requests age out.

Assume every event arrives twice

At-least-once delivery is the only honest guarantee a webhook system can make. Exactly-once sounds better in marketing copy and does not survive a network partition. That means duplicates are the consumer's problem, and your job is to make deduplication trivial: a stable event ID on every payload, documented clearly.

Say the same about ordering. Deliveries can arrive out of order after retries, so consumers should treat each event as a fact about a resource, fetch current state when it matters, and never build a state machine that breaks when step three arrives before step two.

What to look for in a provider

If you're buying instead of building: ask to see the delivery log, ask what happens after the last retry fails, and ask how you replay a dead delivery. Those three answers tell you more about the system than any uptime number, because the happy path already worked in everyone's demo.