Skip to content
ENع
Schedule call
All guides
Integration8 min read

Queues, retries and backfills

AI workloads break queue defaults. Calls take seconds to minutes rather than milliseconds, failures are frequently transient, and cost per attempt is high enough that a careless retry policy is expensive as well as disruptive.

The visibility timeout will bite you

The default visibility timeout on most queues is tens of seconds. A model call with a long context and a slow provider can exceed that, at which point the queue redelivers work that is still in progress and you process it twice.

  1. Set the timeout from your measured p99, with headroom, not from your average.
  2. Better, extend the lease while the work runs, so a genuinely slow job holds its claim and a dead worker releases quickly.
  3. This is exactly the failure that idempotency keys exist for. If you have them, redelivery is a wasted call. If not, it is a double posting.

Retries that do not amplify

The naive policy, retry three times immediately, turns a provider blip into a self-inflicted denial of service at exactly the moment capacity is scarce.

  1. Retry only what is transient: timeouts, 429s, 5xx. Never retry a 400; the input is wrong and will be wrong next time.
  2. Back off exponentially with jitter. Without jitter every worker retries in lockstep and you rebuild the spike.
  3. Cap the attempts, and cap the total time. Work that has been retried for an hour should be somebody’s decision, not the queue’s.
  4. Respect Retry-After when the provider sends one. It is better information than your policy.
  5. Then dead-letter it, with enough context to reprocess without reconstructing anything.

Every retry costs a full model call. A retry policy is a spending policy, and it is worth reading it that way once.

The dead letter queue is a work list

A dead letter queue nobody reads is an outage nobody has noticed yet. Treat it as a queue of decisions rather than a graveyard.

  1. Alert on rate of arrival, not on depth. A steady trickle is normal; a step change is an incident.
  2. Group by failure reason. Forty failures with one cause is one bug, and reading them individually wastes a morning.
  3. Keep the original message and the failure, so replay needs no reconstruction.
  4. Make replay a one-command operation. If it is hard, it will not happen and the work will be quietly abandoned.

Backfills

At some point you will reprocess a month, because a prompt was wrong, a model changed, or a bug corrupted a field. This is where all the earlier discipline pays or does not.

  1. Scope it precisely: a date range, a status, a version. "Everything since March" is not a scope.
  2. Dry run first, writing decisions to the store and nothing downstream. Compare against what is already there.
  3. Read the diff before committing. On one backfill this is where we found that only 3% of records actually changed, which turned a risky rerun into a targeted one.
  4. Run it through the same idempotent path as live traffic. A separate backfill path is a second implementation and it will diverge.
  5. Rate limit it below live traffic so the backfill never starves the queue that is serving people.

Want us to run this with you?

The Audit is this method pointed at your systems, with a costed build plan at the end of it.

Schedule call
Tell us the number you want to move.Schedule call