A failed call does not mean the task has failed
An agent may need to look up information, update a record, and then communicate the result. If a tool response fails midway through that workflow, deciding whether to repeat the last call is not enough. The action may never have reached the external system; it may have executed while its confirmation was lost; or the change may have been applied only in part. Each possibility calls for a different decision.
In this guide, recovery means reconciling the agent’s goal and recorded progress with the state that can be verified in the environment, then choosing a safe way to continue or an explicit stopping point. The unit of analysis is the multi-step task, not an isolated request. Retrying a call can be part of recovery, but it does not define recovery.
The central recommendation is simple: before repeating an action whose result is unknown, check what happened outside the agent. If there is no reliable way to verify it, do not turn uncertainty into a second change. Limit subsequent actions and hand the case to a person when the cost of getting it wrong is greater than the cost of waiting.
Map the failure: first classify what you know
An explicit error does not always prove that the external system was left unchanged. Similarly, a timeout only means that the agent did not receive a response in time; by itself, it does not show whether the operation completed. Classify the incident according to the available evidence, not just the error label shown to the agent.
Distinguish five situations. In a confirmed failure, the response provides evidence that the action did not execute. In an ambiguous outcome, the request may have reached the system, but there is no reliable confirmation. In an invalid response, an answer was received, but its format or content is not safe to use. In an unexpected external state, the system reports something that conflicts with the plan’s assumptions. In an interruption between steps, the task stopped after one or more confirmed effects but before the workflow was complete.
These categories guide the next check; they do not dictate an automatic response. If failure is confirmed and the operation is safe to repeat, another attempt may be reasonable. If the outcome is ambiguous, first look for evidence in the affected system. If the response is invalid or the observed state contradicts the plan, suspend dependent actions until the discrepancy is resolved.
Classification and next step
Use this classification to decide what to verify; it does not replace the guarantees specific to each tool.
| Situation | What is known | Recommended next step |
|---|---|---|
| Confirmed failure | There is evidence that the action did not execute. | Assess whether it can be retried without changing state or duplicating effects. |
| Ambiguous outcome | It is not known whether the action executed. | Check the external system before repeating it. |
| Invalid response | The response is not sufficient to decide or continue. | Validate it or query again; do not treat invalid content as confirmation. |
| Unexpected state | The environment differs from what the plan assumes. | Replan the task or stop dependent actions. |
| Interruption between steps | Earlier effects may be confirmed while later steps remain pending. | Reconstruct progress step by step and resume only from a safe point. |
Before continuing: preserve what is needed to reconstruct the task
A checkpoint is useful when it lets you reconstruct what the agent intended to do and what it knows about each step. You do not need to save all of the agent’s reasoning or every available piece of data. Preserve the minimum operational information: task identifier, goal, relevant constraints, step order, requested tool and operation, parameters needed to identify the operation, response received, confirmation status, and the latest observation of the environment.
Record each step with an explicit status, such as pending, requested, confirmed, failed with evidence, or outcome unknown. Avoid compressing these statuses into a note such as “action completed,” which can erase the distinction between intent and confirmation. If the external system provides a way to look up the modified object, keep the key needed to find it and note when it was last observed.
A checkpoint does not prove that the external world is still the same. It is a snapshot of what the process saved; another person or system may have changed the record between the interruption and resumption. When restoring it, validate again any conditions that matter before making new changes. Microsoft’s workflow checkpoint documentation covers saving and restoring checkpoints; for a specific implementation, the team should check what state is preserved and how it is rehydrated.
The design should also preserve the task’s boundaries: what counts as success, which operations must not be repeated, and which conditions require escalation. Without those boundaries, an agent may reconstruct the technical steps and still continue in a direction that is no longer safe.
Decision tree: resume, verify, replan, compensate, or stop
The decision can be expressed as a short process. First identify the last confirmed step and the first step whose outcome is uncertain. Then ask whether a reliable query can reveal the external state. If so, query before acting. If not, assess the potential impact of repeating the action and whether there is a safe way to resolve the ambiguity. If neither is available, stop and ask for human intervention.
Resuming means continuing from a known point without re-executing steps that are already confirmed. It is appropriate when the saved state is sufficient, relevant conditions still hold, and pending steps are safe. Verifying means querying the environment to find out what happened, not sending the same command again. Replanning means changing the plan because the current state no longer satisfies its assumptions. Compensating means taking a different action to counteract an earlier effect. Stopping means making no further changes until there is a decision or enough evidence.
AWS guidance on checkpoints for agent systems warns that resuming without idempotency guarantees can duplicate effects or corrupt data. This reinforces a practical distinction: saving execution state helps reconstruct the workflow, but does not automatically make repeating an operation safe.
Decision sequence
Apply these steps to the first uncertain point. If an outcome cannot be checked, do not replace it with an assumption.
- 01Identify which steps are confirmed and which is the first without confirmation.
- 02If possible, query the external system for a concrete signal of the expected effect.
- 03If the effect has already occurred, mark the step confirmed and continue only with pending steps.
- 04If it did not occur and repeating it is safe, retry according to the operation’s rules.
- 05If the effect was partial or the state has changed, replan and assess whether compensation is appropriate.
- 06If you cannot verify the state or there is no safe way forward, stop the task and escalate.
Partial effects: undoing does not always mean returning to the previous state
In a composite task, some steps may have completed before a later one failed. If an agent updates a record and then cannot send a notification, repeating the whole workflow could apply the change again, create duplicate records, or send repeated messages. Recovery should start with the confirmed effects and decide what to do with each one separately.
When an operation is reversible, identify in advance what reversal means and how to verify that it took effect. Do not assume that “undo” removes every trace or that the system can return exactly to its previous state. In some workflows, the appropriate response is a compensating action: for example, correcting a record through a new operation instead of deleting the historical operation. That compensation can fail too, so it needs its own confirmation and boundaries.
The saga pattern, described by AWS for multi-step workflows, distinguishes forward recovery—continuing or retrying—from backward recovery through compensating transactions. It is a useful reference for structuring distributed processes, but it does not imply that every agent effect has an available or safe compensation. The choice depends on the system’s rules and the impact of the action.
If an action cannot be reliably reversed, the agent should treat it as a risk boundary. It can record the effect, prevent further steps that would make it worse, and request review. It should not improvise a compensation that has not been defined for the case.
Choosing between continuing and compensating
Use this table as a design guide for each operation with persistent effects.
| Question | If yes | If no or uncertain |
|---|---|---|
| Is the effect confirmed? | Preserve it as part of the progress and assess the pending steps. | Verify the environment before retrying or compensating. |
| Is the next step still valid given the observed state? | Resume from the pending step. | Replan the task; do not follow the old plan out of inertia. |
| Is there a defined, verifiable compensation? | Consider executing it if necessary and authorized. | Do not improvise a reversal; stop and escalate. |
| Is the cost of a duplicate action acceptable and controlled? | A retry may be permissible under the operation’s guarantees. | Require additional verification or human intervention. |
Boundaries and escalation: signals to stop the agent
A recovery policy needs explicit stop conditions. Practical signals include being unable to query the state that would establish whether an action occurred; observing changes that conflict with the plan; receiving repeated invalid responses; accumulating attempts without confirmed progress; exceeding the authorized impact or scope; or lacking a reliable compensation for an unwanted effect. These are proposed design criteria, not automatic safety guarantees.
Also define who receives the case, what information they need, and what actions they can authorize. A useful escalation should include the task goal, confirmed steps, the uncertain point, checks already performed, observed external effects, and the option the agent would have chosen. Avoid presenting a cause as fact if it has not been established.
Stopping does not have to mean silently abandoning the task. The agent can preserve the checkpoint, mark the task as blocked, and clearly communicate what remains to be verified. If the interface offers a resume action, it should check the relevant conditions again rather than assume the environment has remained unchanged.
Test recovery with controlled interruptions
Testing the happy path is not enough, nor is it enough to confirm that a process can restore a checkpoint. Simulate interruptions at different points: before an operation is sent; after it is sent but before a response arrives; after an effect is confirmed but before the next step; and after an unexpected change is observed. Include delayed, invalid, or repeated responses if they could occur in the environment being evaluated.
For each scenario, define the expected behavior in advance: what should be preserved, what query should be made, when it is safe to resume, and what condition requires escalation. Then compare actual behavior with that criterion. The evaluation should include errors of excessive initiative—for example, a duplicate—as well as errors of excessive caution—for example, stopping a task that could have resumed safely.
For monitoring, consider the share of tasks that finish correctly, duplicate effects, inconsistent states, blocked tasks, time to resolve an ambiguity, and the share of escalations that were appropriate. Interpret metrics alongside the severity of each case: a low duplicate rate does not prove that an irreversible operation is safe.
Use the results to adjust checkpoints, verification queries, retry limits, and stop conditions. Keep agent, tool, and external-system failures separate when the evidence allows it; when it does not, record the cause as undetermined instead of forcing an attribution.
Checklist for a recovery drill
Run the test with controlled data and check the agent’s observable behavior, not just whether the workflow can restart.
- 01Choose a multi-step task and identify its external effects.
- 02Mark the points where an interruption could leave an ambiguous or partial result.
- 03Define what evidence would confirm each effect and which actions are reversible.
- 04Interrupt execution at each point and restore the saved state.
- 05Check whether the agent verifies before retrying, resumes only pending steps, and escalates when appropriate.
- 06Record duplicates, inconsistencies, recovered tasks, stopped tasks, and appropriate escalations.
What this guide does not solve
This guide addresses recovery for an agent task that spans multiple steps and may modify external systems. It does not replace API retry design, idempotency guarantees for a specific request, or infrastructure recovery. Those topics still matter: a task protocol can rely on them, but it cannot infer their guarantees when they are undocumented.
Here, idempotency is a property that must be verified for the specific operation before assuming it is safe to repeat; it is not enough to label the entire task idempotent. Likewise, a checkpoint can preserve and restore workflow information, but it does not by itself confirm that the external system applied a change. The thesis on fault-tolerant multi-agent architecture is background work with a different scope: it studies control of a mobile robot and is not a direct recipe for agents connected to business services.
As a final practical check, ask in this order: What is confirmed? What can be verified outside the agent? Which effects have already occurred? Which steps are still valid? Is there a safe compensation? What condition requires stopping? If an essential answer is unknown and acting could make the outcome worse, preserve the state, stop, and escalate.
Open questions
- The provided sources do not establish a universal checkpoint schema or a mandatory set of fields; the proposed minimum data is a design recommendation.
- How to verify whether an operation executed depends on the queries and signals offered by each external system.
- Whether an action is reversible and whether compensation is safe depend on the operation and use-case rules; neither can be inferred from a general pattern alone.
- The sources provide no universal metrics or thresholds for deciding when to retry or escalate; these must be defined and tested for each deployment.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction