Writing webhook consumers for at-least-once delivery
Exactly-once delivery does not exist. How to build webhook consumers that survive retries: idempotency keys from Message-ID, replay-safe effects, dead letters.
Exactly-once is a myth, and that is fine
A producer sends you a webhook. You process it, and the connection drops before your 200 reaches them. The producer now cannot know whether you processed the event, so it must choose: send again and risk a duplicate, or stay silent and risk a loss. There is no third option, and no amount of queue vendor marketing creates one.
At-most-once loses events, which is rarely acceptable for email-driven workflows: a dropped order confirmation is an order you never shipped. So every serious producer picks at-least-once, retries until acknowledged, and hands the remaining problem to you. Duplicates will arrive, and your consumer must make them harmless.
The honest framing is exactly-once effect built on at-least-once delivery. Delivery is the producer's half. Effect is yours, and it is entirely achievable.
Where duplicates actually come from
Duplicates are not a freak event; they are scheduled. The classic case is the slow 200: you processed the event in nine seconds, the producer's timeout was eight, and the retry is already in flight carrying the same payload.
Deploys interrupt handlers mid-request. Operators press the replay button while debugging, which is a feature, not an accident. The producer recovers from an outage and re-delivers a backlog. And in email pipelines, the origin sender sometimes re-sends the same message, a duplicate born upstream of any webhook.
Design for all of these at once and none of them stays interesting. That design is one table and one habit.
Idempotency keys: Message-ID is already there
An idempotency key names the event, never the delivery attempt. For email-driven webhooks the natural key already exists: the Message-ID header, unique per message by convention and present in your payload. The payload's own id works the same way. What you must never use is anything tied to arrival, like a timestamp or a per-request uuid.
Processing then becomes claim first, work second: insert the key into a table with a unique constraint, and if the insert affected zero rows, someone already claimed this event, so acknowledge and stop.
On HideMy.world part of this is handled before you: duplicates from the origin sender are removed upstream by Message-ID, so a re-sent email reaches your endpoint once. The claim table still earns its place, because delivery retries of the same payload are a different animal, and those are yours to absorb.
-- once, in your schema
CREATE TABLE processed_events (
event_id text PRIMARY KEY,
seen_at timestamptz NOT NULL DEFAULT now()
);
// per delivery, inside the handler
const claimed = await db.query(
"INSERT INTO processed_events (event_id) VALUES ($1) ON CONFLICT DO NOTHING",
[payload.headers["message-id"]]
);
if (claimed.rowCount === 0) {
return res.status(200).end(); // duplicate: acknowledge, do nothing
}
await enqueue(payload); // first and only timeReplay-safe side effects
The claim table protects you only if the work behind it is replay-safe too, because a crash between claim and effect still needs a recovery story. Sort your side effects into three bins.
Naturally idempotent effects can simply run again: an upsert keyed on order number, setting a status to shipped, writing a row keyed by the event id. Prefer these shapes wherever you have a choice; they turn the whole question into a non-event.
Effects that are idempotent with help need the key passed along: many APIs accept an idempotency key header, so forward yours. Inherently unsafe effects (increment a counter, append a line, ping a human) must sit strictly behind the claim, and where your database allows it, claim and effect belong in one transaction. For external calls, record intent in that same transaction and execute from the record, so a crash replays from your own table rather than from the producer's retry.
Respond fast, let the producer back off
Acknowledge when the payload is safely stored, not when the work is finished. A handler that parses PDFs inline will blow through the producer's timeout, collect a retry, and manufacture the exact duplicate problem it was supposed to avoid. Persist or enqueue, return 200, process in the background.
When you genuinely fail, fail honestly with a 5xx and let the producer's backoff do its job. A well-behaved producer retries with exponential backoff and keeps a log of every attempt; HideMy.world's webhook channel does exactly that, and lets you replay a delivery manually once your fix ships. Verify authenticity before any of this: each delivery is signed with HMAC-SHA256 over the timestamp and raw body, sent in the X-HideMy-Signature header, so forged events die at the door.
The division of labor is clean. The producer owes you retries, per-attempt logs, and replay. You owe it fast acknowledgments and honest status codes.
Poison messages
A poison message fails on every attempt: a payload that trips a bug in your parser, a record that violates one of your constraints. Retries cannot fix it, so backoff just stretches the same crash over hours and buries healthy deliveries behind it.
The pragmatic defense is a dead letter path. Catch the failure, store the raw payload with the error, return 200 so the producer stops retrying, and alert on the dead letter table. You have converted an infinite retry loop into a finite to-do item.
Then fix the bug and replay: from your own dead letter store, or from the producer's delivery logs when it offers manual replay. Cap attempts on your side as well; anything that dead-letters three times deserves a human before a fourth try.
Keep the dead letter table visible rather than buried. A small count on an internal dashboard or a daily digest message is enough; the failure mode to avoid is a dead letter table nobody has opened since March.
A consumer checklist
None of this is heavy machinery. The claim table is ten lines, the dead letter path twenty, and together they absorb every delivery pathology a producer can throw at you.
Exactly-once will stay a myth. A consumer built along these lines makes it an irrelevant one.
- Verify the HMAC signature on the raw body before touching the payload
- Claim the idempotency key (Message-ID or payload id) with a unique insert
- Acknowledge after persisting, do the real work in the background
- Make effects upserts where possible; pass idempotency keys downstream where not
- Return 5xx for transient failures; dead-letter and return 200 for permanent ones
- Alert on the dead letter table and on sustained retry streaks
- Keep a manual replay path and rehearse it before you need it