Where the duplicate comes from
The dangerous window is between the downstream system committing and your system recording that it committed. Everything in that gap looks identical to a failure, and the safe-looking response, retrying, is the one that double-posts.
- The worker is terminated after the POST succeeds but before the response is read. Autoscaling does this routinely.
- The queue redelivers because the visibility timeout expired while the call was slow.
- The downstream returns a 500 having already applied the change, which is more common than its documentation admits.
- A well-meaning operator reruns a failed batch.
The key is the whole design
Everything rests on a key that is derived from the work, not from the attempt. Generated per attempt, it is useless.
- Good
- A hash of the source document id, the target record and the operation. The same work always produces the same key, from any worker, on any retry.
- Bad
- A UUID minted when the job starts. Every retry is a new key and every retry is a new posting.
- Also bad
- A timestamp, a queue message id, or anything that changes when the message is redelivered.
Put the key in the ingress component, where work enters, not in the step that writes. By the time you are three steps deep it is too late to derive it from anything stable.
Three ways to enforce it
- Best: the downstream accepts an idempotency key and does the work for you. Many modern APIs do. Send it and record what came back.
- Next: a unique constraint in the target system on a field you control. The database rejects the second write, and a constraint violation is a success, not an error. Handle it as one.
- Last: your own ledger. Before writing, insert the key into a table with a unique index; if the insert fails, the work is already done. Commit the ledger row and the write in the same transaction where you can, and where you cannot, write the ledger first and reconcile.
The third option has a real gap: ledger written, downstream write fails. You are now marked done and are not. This is why the reconciliation job in the next section is not optional.
Reconcile, because the gap is real
Whatever you do, some small number of records will end up in a state where your system and the downstream disagree. A daily job that finds them costs an afternoon and is the difference between a discrepancy you fix and one a customer finds.
- Compare your decision store against the downstream for the last 48 hours, keyed on the idempotency key.
- Report both directions: marked done but absent downstream, and present downstream but not recorded here.
- Fix the first automatically by retrying; the key makes that safe. Escalate the second to a person, because it usually means a write happened that you did not initiate.
- Alert on the count, not on individual rows. A steady one or two a day is the system working. Forty is an incident.
Every write-back we run in production has this job. It has caught a queue misconfiguration, a downstream silently rejecting a field, and one case where a scheduled job in another team was writing to the same records.
Want us to run this with you?
The Audit is this method pointed at your systems, with a costed build plan at the end of it.
Schedule call
