Home Backend

How to Design Idempotent REST APIs for UK Payment Webhooks

Backend

September 20, 2026

How to Design Idempotent REST APIs for UK Payment Webhooks

A duplicate webhook is rarely a technical curiosity. It is a customer seeing two top-ups, a finance team unpicking a refund that was applied twice, or a support ticket that eats an afternoon. UK payment rails make this more likely than you might expect: Faster Payments settles in seconds, Direct Debit runs across the Bacs cycle over several working days, and providers can only tell you what happened by retrying a message, sometimes hours later. Idempotency is how you make those retries boring.

Why the same event turns up twice

Providers retry because your 500 tells them nothing about what actually happened. A request that records a payment and then dies before writing a response will be sent again, and it should be: the provider cannot know whether the work completed. Meanwhile, two workers can consume the same queue message, a deploy can replay a dead-letter queue, and someone will always re-send yesterday's events after a database restore.

None of that is a provider bug. It is the deal you accept with HTTP and at-least-once delivery. Your endpoint therefore needs to be safe to call many times with identical content, and change state exactly once.

The event ID is your idempotency key

Every serious provider gives each event a unique identifier. Stripe puts one on the event object; most others do the equivalent in their own shapes. That identifier, not the payload body, is your idempotency key. Hashing the JSON is tempting but brittle. Key order and whitespace can vary, providers add fields without warning, and you end up deduplicating deliveries that were never the same thing.

Two details matter. First, verify the signature against the raw request body before you parse or trust anything; a 401 is far cheaper than a fraudulent refund. Second, an event can carry more than one resource, so key on the event id for the delivery and on a business key, such as payment plus effect, for the state change itself.

The same principle applies outbound. When you call a provider to create a payment or a refund, send an idempotency key header, so a timeout on your side does not become two charges because the first request actually succeeded.

Where deduplication belongs

Do not deduplicate in memory. A dictionary inside your application process vanishes on the next deploy and is not shared between instances. Use a table with a unique constraint and let the database arbitrate. It is the only component that can settle the race between two workers holding the same event.

A pattern that holds up under real traffic:

  1. Verify the signature against the raw body and reject anything that fails.
  2. Insert into a webhook_events table with a unique index on (provider, event_id).
  3. If the insert conflicts, return 200 and stop. You have seen this event before.
  4. Otherwise write the event row and enqueue a job in one transaction, then acknowledge immediately.
  5. Mark the row as processed when the job finishes, and run a sweeper that re-queues rows stuck in a processing state for more than a few minutes.

Keep the raw payload for a short retention window. It makes replay and audit possible, and finance queries about what a provider actually sent on the fourteenth stop being guesswork. Treat that payload as personal data where it identifies a customer, and give it a deletion date.

Make the effect idempotent too

Deduplication by event id protects you from the same delivery arriving twice. It will not save you from two different events describing the same outcome, or from a bug that triggers the same effect along two code paths. Put a constraint where the money is: a unique index on the refund or ledger entry, or a status check that only permits pending to succeeded once. If a duplicate slips past your handler, the database should refuse the write rather than quietly double a balance.

Out-of-order events and UK settlement cycles

Webhooks arrive in whatever order the network decides. A refund notification can land before the payment it refers to, and a retry of an old event can arrive after a newer one. Two habits fix most of this:

  • Store the event's own timestamp or sequence number, and ignore anything older than what you have already applied for that resource.
  • Model each payment as a small state machine, for example created, authorised, captured, settled, refunded, chargeback, and reject transitions that make no sense from the current state.

The UK context makes this concrete. Faster Payments moves money in seconds, but the notification telling you about it can lag, so a payment that looks settled in your dashboard at 09:00 may still be pending in a report generated at 08:59. Direct Debit runs on the Bacs cycle, where a payment can sit pending for several working days before it collects or fails. Give your states room for that dwell time, and treat webhooks as a fast notification rather than the final word. Reconcile against the settlement or payout reports your provider supplies, particularly before you release goods or mark an invoice as paid.

What to return, and when to let it fail

The acknowledgement is part of the contract, so get it right:

  • 2xx for handled events and for duplicates. A duplicate is success, not an error.
  • 4xx only when a retry cannot possibly help, such as a failed signature, a malformed body, or an account you do not recognise.
  • 5xx when you want the provider to try again, for example when your database or queue is unavailable.

Acknowledge quickly and do the work asynchronously. Watch queue depth, and give failed jobs a dead-letter queue with an alert attached. Returning 200 and then dropping the message silently is the worst of both worlds: the provider stops retrying and you never find out.

Test idempotency before production

Write a test that posts the same event twice and asserts a single ledger entry. Then post it twenty times concurrently and assert the same. Add a test where events arrive out of order, and one where the handler throws halfway through, so you can confirm the sweeper picks the work back up.

In staging, use the provider's test mode and replay a captured payload. If their dashboard lets you resend an event, do it, because there is no substitute for a genuine retry with real headers and a real signature. Then check the thing nobody checks: your logs should show two deliveries and one business effect, clearly distinguished from each other.

Ship the boring guarantees

Idempotency is not clever architecture. It is a handful of constraints plus a rule that work never happens twice, no matter how many times the message arrives. If you take one thing from this, take the unique index. It is a few lines of migration, it costs almost nothing at runtime, and it will be sitting there at 3am when a provider retries a day of events after an outage.

If your payment flows touch invoicing, VAT or financial reporting, check the specific requirements with your accountant or compliance adviser before you rely on webhook state alone.

Photo: tomekwalecki / Pixabay

Related Posts

Developer Laptop Setup Checklist for New UK Hires
Tools

October 10, 2026

Developer Laptop Setup Checklist for New UK Hires

A practical checklist for setting up a secure, comfortable development laptop as a new UK hire, from disk encryption and access requests...

read more
A Beginner's Guide to Database Normalisation for Small Business Apps
Databases

October 09, 2026

A Beginner's Guide to Database Normalisation for Small Business Apps

A practical introduction to first, second and third normal forms, with clear examples showing how to structure small business data...

read more
How to Run Zero-Downtime Database Migrations
Databases

October 07, 2026

How to Run Zero-Downtime Database Migrations

Practical steps for changing production schemas without downtime: the expand-and-contract pattern, lock-aware statements, deploy...

read more