The problem: a lost response does not mean an operation failed
In an integration with an AI API, a timeout usually describes a limited fact: the client did not receive a response within the configured deadline. On its own, it does not establish that the provider did not receive the request, did not process it, or did not initiate a subsequent action. The connection may break after the server accepts the request; the response may be lost after the operation completes; and a local process may restart when the result already exists outside that process.
This distinction matters even more when a model output triggers operational effects. Repeating a text generation may produce a different response and consume resources, but it will usually not change a business system. Repeating an instruction that sends an email, creates a booking, records a payment, modifies a case file, or invokes a tool can produce a duplicate effect. The safety of the flow must not depend on the model returning the same text or on a single network call always reaching a successful conclusion.
HTTP semantics help define the boundary of the problem: an idempotent request is one whose intended effect on the server remains equivalent even if it is executed more than once. This does not mean that every repetition is free, that the response will be identical, or that no additional records will be created. It means that the relevant effect must be applicable repeatedly without changing the expected final result. In AI flows, that property must be verified both for the provider call and, independently, for every external system receiving an action.
A mental model: request, logical operation, attempt, effect, and confirmation
It is useful to separate five concepts that are often mixed together. The logical operation is the business intent: for example, “produce a structured response for case X” or “send one approval notification.” A request is a specific message sent to an API. An attempt is each transmission of that request, including the first send and its retries. The external effect is the observable change at a destination: a sent message, a created row, or a confirmed purchase. Confirmation is the evidence that allows the logical operation to be marked as completed, rejected, or awaiting review.
A single operation identifier should connect every attempt pursuing one business intent. It must be created before the first call and persisted outside process memory so it survives failures, deployments, and background jobs. Do not use only a request identifier returned by a provider as the identity: it may not exist after an early failure and, even when it does, it will normally identify an attempt rather than the full business intent.
In addition to the operation identifier, retain a stable idempotency key for any receiver that supports one. The key must remain unchanged when repeating the same logical operation and change when the intent changes. Generating a new key on every retry defeats deduplication. Reusing one key for two different operations may cause a legitimate action to be mistaken for a repeated one. The record should retain the relationship among the operation, attempt, destination, key used, and observed result.
Identities that must not be confused
| Item | Scope | Design rule |
|---|---|---|
| Operation identifier | Business intent | Created once and retained until the operation is closed. |
| Attempt identifier | One specific transmission | Changes for every call or retry. |
| Idempotency key | Contract with a receiver | Remains stable for the same operation at that receiver. |
| Effect identifier | Destination system | Stored when the destination confirms the change or can locate it. |
Classify risk before automating a retry
No single policy is appropriate for every call. Classifying an operation by its repetition risk forces you to decide what must be protected. Reads or queries without effects can usually be retried when cost and load are controlled. Text generation without external effects can be repeated, but the result may vary; the application must therefore decide whether it accepts a new output, keeps a partial one, or presents the case as pending.
An idempotent write can be repeatable if the destination guarantees that the same key represents the same operation. A compensable write may require a later reversal, but compensation does not automatically make a retry safe: it can also fail, arrive late, or produce effects of its own. Irreversible or difficult-to-verify actions require a stronger barrier, such as human confirmation, a prior reservation, or a reliable query to the destination system before acting.
This classification must be applied to the complete flow, not only to the model call. A model may correctly generate a tool call, while the tool itself may have executed before the response was lost. The subsequent text output is not enough evidence that the effect occurred exactly once. The layer that runs tools must record and deduplicate the action with its own controls.
Decision matrix for retries
| Operation type | Risk of repetition | Initial policy | Closing evidence |
|---|---|---|---|
| Text generation without effects | Alternative result and additional consumption | Bounded retry if the deadline permits | Stored response or final error state. |
| Structured output | Incomplete data or invalid format | Correct validation or repeat according to the contract; do not assume identical content | Validated schema and stored result version. |
| Read tool call | Additional load or changing data | Retry with concurrency limits | Tool response and timestamp. |
| Idempotent write | Duplicate if deduplication is weak | Retry only with a stable key and persistent record | Confirmation or lookup of the created resource. |
| Irreversible external action | Double effect or irreparable effect | Do not retry automatically after an ambiguous state | Unambiguous confirmation or human review. |
Design idempotency at two boundaries
Deduplicating a provider call and deduplicating a business effect are related problems, but they are not equivalent. Even if a model API accepts an idempotency key, that protection does not prove that a downstream tool, email provider, or payment system applied its effect only once. Likewise, an idempotent tool does not eliminate the cost or saturation caused by unnecessarily repeating an inference request.
The most robust practice is to establish two boundaries. At the first boundary, the API wrapper records the operation and associates attempts with a stable key when the provider contract supports it. At the second, the external-action executor uses its own effect identifier and durable deduplication store. Before execution, it checks whether a completed action already exists for the operation; if one does, it returns the existing result. If not, it records the start in a way that allows a later restart to continue the investigation.
Do not invent reconciliation capabilities that the actual contract does not provide. The available sources describe general retry practices and HTTP semantics, but they do not document a universal operation-status lookup or an idempotency key applicable to every AI API endpoint. Review the contractual documentation for the specific endpoint before depending on any such capability.
Minimum flow for an operation with an effect
- 01Create and persist the operation identifier before making remote calls.
- 02Record the initial state, intent, destination, and version of relevant data.
- 03Send the attempt with the same idempotency key when the receiver supports it.
- 04If a valid confirmation arrives, store the result or effect identifier and close the operation.
- 05If a timeout or disconnection occurs, mark the state as ambiguous; do not create a new operation.
- 06Query available status or reconcile against the destination using the stored evidence.
- 07Retry only if the policy for the operation type permits it; otherwise route the case for review.
Retry policy: budget, backoff, and jitter
A safe policy expresses limits before an error occurs. It should define which error families are candidates for retry, the maximum number of attempts, a global operation deadline, a maximum acceptable wait, and the maximum tolerable cost. The number of attempts alone is not enough: five retries may exceed the user deadline, exhaust a quota, or keep workers occupied when they should release capacity.
Rate-limit responses and transient server errors can justify waiting and retrying, provided the operation is repeatable and budget remains. OpenClaw documentation states that certain Stainless-based SDKs can treat 408, 409, 429, and 5xx-family responses as retryable. That describes an SDK policy in that context, not a universal rule for every endpoint or external effect. Errors indicating an invalid request, missing authorization, or a business condition are not fixed by repeating identical data; they require correcting the cause or stopping the flow.
Use exponential backoff to progressively extend the interval and jitter so that many clients do not retry at the same time. AWS recommends both exponential backoff and random variation, as well as limiting retries and verifying idempotency before repeating. If a response indicates how long to wait through a retry signal, honor it when valid and compatible with the operation deadline. If that wait exceeds the budget, record the reason for deferral or failure instead of waiting indefinitely.
Rate limits and saturation: a retry is also load
A 429 error means that available capacity or an applicable limit does not allow progress at that time; it does not prove that applying more pressure will solve the problem. Retrying immediately can turn a contained incident into a traffic storm. In addition, failed attempts may count toward rate limits, so an aggressive strategy can delay valid operations even further.
Control concurrency in the work queue, not only inside each client. Set limits by provider, model, credential, and operation type where appropriate. Reserve capacity for reconciling ambiguous states and for high-priority operations; otherwise, a wave of retries can prevent the system from determining what happened. The budget should include queue time, connection time, processing time, and waits between attempts.
OpenAI guidance on 429 errors recommends exponential backoff with jitter when no usable wait indication is available, and advises limiting both the number of retries and the total time spent on them. It is important to distinguish transient throttling from other account or quota problems that waiting does not fix. The decision must be based on error information and the integration contract, not on the status code alone.
Ambiguous states: reconcile before repeating
An ambiguous state appears when there is not enough confirmation to determine whether the effect occurred. It must be an explicit, persistent state, not an exception erased when a process restarts. Record at least the operation identifier, the data or a safe summary of the intent, attempt identifiers, timestamps, the idempotency key, the destination, the error category, and any identifier returned before the interruption.
Reconciliation follows a hierarchy. First, use a status query or resource identifier if the receiver contract provides one. Next, look for the effect in the destination system using a stable criterion, such as the operation identifier included in metadata. If evidence confirms the effect, close the operation without repeating it. If it proves the effect was not applied, you may open a new attempt under the policy. If it cannot distinguish the two cases, do not assume absence: keep the case pending and escalate it when the risk warrants it.
Human review is not a design failure; it is a safety control for operations whose cost of duplication exceeds the cost of delay. Escalation conditions should be concrete: financial charges, irreversible communications, modification of regulated records, inconsistency between sources, expiration of the deadline, or absence of reliable proof of deduplication.
Decision after a timeout following transmission
- 01Mark the attempt as having an unknown response and retain all available evidence.
- 02Check whether the receiver provides a status query, retrieval by key, or resource identifier.
- 03Check the system that received the effect, not only the model or agent layer.
- 04Close as completed if sufficient evidence of the expected effect exists.
- 05Retry only if there is evidence of non-execution or an applicable idempotency guarantee.
- 06Escalate if evidence remains ambiguous and the effect could be consequential or irreversible.
Patterns for structured outputs, tools, and agents
For structured outputs, separate validation from execution. A response that does not satisfy the schema must not directly feed a tool. Store the received output, validate types, required fields, value ranges, and authorization for the proposed action. If you decide to request a new generation, treat it as a new attempt to produce a plan, not as proof that no tool ran earlier.
In tool calling, the controller must be the execution authority. The model may propose a call, but the controller must assign the operation identifier, verify permissions, deduplicate semantically equivalent arguments when appropriate, and record the result. If the agent resumes after a failure, it must recover the record of tools already executed; it must not infer history from the text of a conversation.
For agents with multiple steps, avoid retrying the entire composed flow as an opaque unit. Retry individual steps only when their boundaries and guarantees are known. Planning can be regenerated; a read can be repeated with a warning that data may have changed; a write must be reconciled; and an irreversible action requires an explicit barrier. This design reduces duplicates and also improves auditability when the system receives partial results.
Production checklist and limits of this guide
Before enabling automatic retries, document the owner, destination, effects, duplication cost, idempotency key, confirmation evidence, deadline, maximum number of attempts, and escalation condition for each operation. Test failures at every boundary: before sending, during transmission, after the destination accepts the request, and before persisting the local response. A useful test verifies that restarting a process does not create a second effect.
Also review the architecture in the context of your product’s learning center, safety documentation, pricing criteria, and glossary. Retry limits affect cost and latency; record retention affects privacy; and credentials used to reconcile or run tools should have least privilege. For specific models, such as GPT-6 Astra, and for integrations with an organization such as OpenAI, the final policy must be adapted to the contract and capabilities actually documented for the endpoint in use.
The closing rule is simple: do not declare success because you emitted a request, and do not declare that an effect is absent because you did not receive a response. Declare an operation resolved only when persisted evidence supports its state. When that evidence is unavailable, the safe decision may be to wait, reconcile, or request human intervention.
Checklist before allowing an automatic retry
| Question | Required answer |
|---|---|
| Does the logical operation have a persistent identifier? | Yes, created before the first transmission. |
| Does the receiver support verifiable deduplication? | Yes, through a documented key or query; otherwise reconciliation is planned. |
| Does the external effect have its own protection? | Yes, independent of the model call. |
| Is there an attempt limit and global deadline? | Yes, with a time and cost budget. |
| Do ambiguous states have handling? | Yes, with defined evidence, query, and escalation. |
| Has a restart between acceptance and response been tested? | Yes, and it does not produce a second effect. |
Open questions
- The supplied sources do not document an idempotency key or status query that is universal across all AI APIs; these capabilities must be verified for the specific endpoint.
- Retryable error categories may vary by provider, SDK, endpoint, credential, and operation type.
- A 429 response can arise from different limits or account conditions; the code alone does not determine the correct action.
- The ability to locate an external effect after a timeout depends on the destination system retaining and allowing lookup of a stable identifier.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction