Design retries that do not repeat business actions
A repeated request is not always a harmless second try.
Understand the ambiguous outcome
A request can complete on the remote service while its response is lost. Retrying a payment, email or record creation without checking can repeat the effect. Record a stable operation identifier and use the provider’s supported idempotency mechanism where available.
Check the provider’s retry contract
Keep the same key and payload for retries of one logical operation, within the provider’s documented retention rules. Stripe, for example, compares request parameters and can return the original result even when it was an error. Once an earlier key record has been removed, reusing that key can create a new request. These are provider-specific rules: a key is not a permanent guarantee against duplicates. Do not generate a new key merely to escape an ambiguous result; reconcile it first.
Respect service feedback
Rate-limited services may return HTTP 429 and a Retry-After value. Retry-After can be an HTTP date or a delay in seconds, so check the format before scheduling a retry. Follow documented limits instead of immediately repeating requests. Use bounded retry counts and an overall deadline; retrying forever can turn a small failure into an expensive queue.
Store enough evidence to reconcile
Keep the intended operation, request identifier, attempt times and confirmed remote result. Avoid logging passwords, API tokens or unnecessary personal data. If the outcome is unknown, query the remote state using the documented identifier before issuing another write.
Test recovery with synthetic records
Simulate a timeout before the response, a duplicate delivery and a worker restart. Confirm that the intended effect occurs once and the run reaches a clear terminal state. If the provider offers no safe deduplication or lookup, require human reconciliation for ambiguous writes.
Keep working through the question
Automated Empire · Published . Updated . Prepared with AI assistance; editorial approach and corrections are described on our method page. Examples are illustrative, not case-study results.
References and further reading
- Stripe: Idempotent requests — A provider-specific example of matching parameters, cached results and expiring key records; other APIs can differ.
- RFC 9110: Retry-After — Defines HTTP-date and delay-seconds forms of the Retry-After response field.
- MDN: HTTP 429 Too Many Requests — Explains rate limiting and the optional Retry-After response header.