Skip to content

Engineering · 10 September 2026 · 11 min read

The idempotency contract for wallet writes

Most stranded balances trace back to five decisions. Here is the contract we implement, including the one clause that teams consistently get backwards.

Every platform that moves player money needs an answer to one question: what happens when the same operation arrives twice. Payment providers resend callbacks. Players retry on a slow network. A deployment restarts a process mid-write. The answer cannot be "it depends", and it cannot live in each caller.

Clause one: the client fails closed

The shared helper that writes a balance must throw when no idempotency key is supplied. The temptation is to generate one as a fallback so nothing breaks. That fallback is the bug: a generated key is unique per attempt, so a retry produces a new key and the operation applies twice. Failing closed is noisy in development and correct in production.

Clause two: keys are derived, never random

The key has to be a deterministic function of the operation, so the same logical action produces the same key on every attempt. A payout reserve keys on the payout reference. A refund keys on the order reference. An administrative credit keys on the user and a time window. If you cannot write the key down before making the call, the operation is not yet defined well enough to be safe.

Clause three: check and insert atomically

Checking for an existing key and then inserting one is a race with a window in it, and under load something will land in that window. The check and the insert have to be a single atomic upsert that tells you whether you created the record or found it. Then the branches are explicit: found and still processing means reject with a retry-after, found and completed means replay the stored response, and a duplicate-key error from a concurrent upsert takes the same path as processing.

Clause four: persist completion before you answer

When the handler finishes, the status and the response body must be written before the response ships. Fire and forget leaves a window where a retry arrives, sees a record still marked processing, and gets refused for an operation that actually succeeded. It is a small await and it removes an entire class of support ticket.

Clause five: cache a success, reclaim a refusal

This is the one teams get backwards, and the consequences are quiet. Because the key is deterministic, whatever you store against it is served to every later attempt at that operation. So the outcomes have to be split.

  • A success is cached with its body and status. A retry replays it, which is the entire point.
  • A client error is stored as failed with no body, and the read path reclaims that record so the caller can try again. Guard the reclaim on the status, so of two racing retries the loser matches nothing and is refused rather than both proceeding.
  • A server error is cached deliberately. This is the opposite call from a client error, and it is intentional: a failure inside your own system may have applied a partial side effect, and re-running it under the same key is exactly the double-apply that idempotency exists to prevent. Recovery there is an operator decision, not an automatic retry.

Why it matters

Without the reclaim rule, a caller refused for a reason they then fixed is served the cached refusal forever and can never succeed. Without the server-error rule, a partially applied operation gets replayed on the next retry. Both failure modes are silent, and both are expensive.

The part that is not code

Write the contract down once, in one place, and make every copy of the middleware point at it. We have audited estates where seven copies of the same middleware had drifted, five of them on the clause above, purely because the reasoning had never been recorded next to the implementation. Drift here does not surface as a failing test. It surfaces as money in the wrong place.

Related service

Software Architecture

Architecture that still holds when the second brand, second product and first regulator arrive.

Read the service page

More insights

Next step

Tell us what you are building, or what you are about to buy.

One working day to a reply, from an engineer rather than an account manager. Under NDA as standard, before anything is shared.