From chatbot to agent: situated autonomy, not blind trust
An agent is not simply an assistant that responds in a conversation. For design purposes, a useful operational definition combines a model that interprets a task with execution state, access to tools, a policy that defines permitted decisions, and a loop that observes outcomes and selects the next step. It may query a database, open a ticket, look up information, update a record, or invoke an API. The important distinction is not that the model “reasons,” but that the system can affect systems beyond the chat window.
This distinction changes the starting question. Do not begin with “Which model performs best?” Begin with “What can this system do, under which identity, against which resources, and under what conditions?” Model capability can help complete a task, but it does not replace authorization boundaries, result validation, or oversight. The same conclusion applies when reading analyses of GPT-6 Astra, Claude Opus 5, or Gemini 3.8 Flash: higher claimed or measured capability does not remove the need for operational controls.
Not every task requires an autonomous agent. A fixed flow with predictable steps and few exceptions may be better handled through conventional automation or a workflow in which the model only classifies or drafts a proposal. Added autonomy must be justified by task variability and by the ability to contain its effects. As a starting rule, an external action should be more restricted than a query, and a change that is difficult to undo should require more evidence and more oversight than a reversible action.
The recommended entry card for the hub page, [Agentes: diseño, evaluación y supervisión](/learn), should present this guide as an architecture and decision framework, not as a universal recipe or a vendor comparison.
Before connecting a tool: classify the impact of every action
A tool should not be enabled merely because it looks useful in a demonstration. It needs an action record: purpose, affected systems, data it receives, identity used, permitted operations, volume limits, whether its effect is reversible or irreversible, external dependencies, and human owner. This record makes a commonly hidden distinction visible: “look up orders” and “cancel orders” may use the same API, but they do not carry the same risk.
Classification can combine four dimensions. Impact estimates the harm if the action is wrong; reversibility determines whether it can be safely undone; scope measures how many accounts, records, or systems may be affected; and sensitivity assesses the data or secrets exposed. A fifth practical dimension is ambiguity: if a request permits several reasonable interpretations, the action should not be performed autonomously even where that is technically possible.
External communications deserve their own category. Sending an email to a customer, publishing a response, or creating an incident may be technically reversible, but not necessarily reversible in reputational, contractual, or privacy terms. The same applies to production changes: having a rollback operation does not make an incorrect deployment harmless. For these cases, the design should provide for review of the content, recipient, scope, and execution time.
Early in implementation, connect this classification to the guidance on [irreversible actions and approvals](/safety/acciones-externas-agentes) when it is available. The goal is for approval not to be a generic confirmation gesture, but an informed decision about a specific action, its parameters, and its possible consequences.
Initial autonomy matrix by action type
| Situation | Example | Recommended initial autonomy | Minimum control |
|---|---|---|---|
| Low-impact query | Read the status of a request | Read-only | Resource filtering and access logging |
| Bounded, reversible change | Update a non-critical field | Limited execution | Parameter validation, scope limit, and rollback option |
| External action with meaningful effect | Send a communication to a customer | Proposal with approval | Preview, explicit recipient, and human approval |
| Production change or deletion | Modify configuration or delete data | Not autonomous | Strengthened approval, checkpoint, and recovery plan |
Least-privilege permissions: controls that do not depend on the model’s text
Least privilege means granting only the permissions needed for a defined task, for the time needed, and against the resources needed. For agents, this requires more than a broad token with extensive access. A credential that can read, edit, and delete every resource turns an incorrect output, manipulated instruction, or integration failure into a far-reaching incident.
A sound practice is to separate identities by tool, environment, and purpose. A support agent should not use the same identity as a deployment process, and a test identity should not access production. Short-lived credentials reduce the exposure window, but they do not by themselves correct excessive authorization: specific access scopes, quotas, network restrictions, and server-side validation are also required.
The tool must verify authorization deterministically. An interface exposing concrete operations—for example, creating a reply draft for an assigned case—is preferable to a generic interface that permits arbitrary queries or commands. Where the system uses allowlists, maximum amounts, or roles, those limits must be enforced outside the context the model can alter. Prompt instructions can guide behavior; they are not a sufficient security boundary.
Connectors should treat all content received from web pages, documents, emails, tickets, and tool responses as potentially untrusted data. The presence of an instruction in text does not give it authority. The action policy, service identity, and parameter validator must take precedence over any instruction encountered while carrying out the task.
This section should link to [sensitive data and credentials](/safety/datos-sensibles-agentes) when available. For teams handling personal data, minimization also means avoiding the delivery of more attributes to a tool than are needed to complete an action and defining who may access the resulting records.
Tool onboarding process
- 01Define a specific operation, its owner, and its expected outcome.
- 02Document the strictly necessary resources, data fields, environments, and operations.
- 03Create a separate identity with narrow-scope, short-lived permissions.
- 04Implement schema validation, volume limits, and authorization in the receiving service.
- 05Test denials: unassigned resources, out-of-range parameters, and calls from the wrong environment.
- 06Enable logs and a revocation mechanism before making the tool available to the agent.
When to require a human in the loop
Human review is most useful when it is integrated at a defined decision point. Requiring confirmation for every low-risk query slows work and encourages mechanical approvals; not requiring it for meaningful actions shifts the burden of detecting errors to those who experience their effects. The design should establish explicit thresholds and give the reviewer the information needed to decide.
An approval should display the planned action, final parameters, affected systems, the identity that will execute it, the agent-generated justification, and the effect of approving or rejecting it. If the action depends on uncertain facts, the interface should also show that uncertainty rather than present a recommendation as though it were a verified finding. Approving an abstract intent, such as “resolve the case,” is less safe than approving a bounded operation, such as “send this draft to this recipient.”
As a starting point, deletions, payments or financial commitments, production changes, external communications, permission changes, processing of especially sensitive categories, and operations with ambiguous requests should pass through human review. The organization may add thresholds for amount, number of records, or operational severity. These thresholds are internal policy decisions: no universal figure makes an action safe.
Human oversight must not be treated as an excuse to neglect the system. The reviewer needs genuine authority to reject, correct, and escalate; training on the process; and a workload that makes review possible. If they receive hundreds of nearly identical requests, the control can become a formality.
Useful memory without indefinite retention
Memory can improve task continuity, but it also increases the privacy surface, the risk of using stale information, and the difficulty of correcting errors. At minimum, distinguish between session context, operational task state, authorized persistent preferences, knowledge retrieved from documentary sources, and audit records. These categories serve different purposes and should not automatically share the same retention period or permissions.
Session context maintains coherence during an interaction and will usually expire at its end or after a short defined period. Operational state retains information needed to resume work, such as a case reference or checkpoint. Persistent preferences require a clear purpose, known provenance, and a mechanism for access and correction. Retrieved knowledge should retain its source, version, and date so the system can identify when the source has been replaced or invalidated.
Provenance matters as much as content. If a memory comes from a user, a business system, or an inference made by the agent itself, the system should distinguish those origins. An unverified inference—for example, an inferred priority or preference—must not be treated as confirmed data. Persistent memory should also not become a route for keeping secrets, irrelevant data, or instructions injected through external content.
Before storing anything, answer four questions: which specific purpose it serves, who can read it, when it expires, and how it is invalidated. The future guide on [memory, retention, and expiry](/learn/memoria-agentes-caducidad) can develop these policies further. In the European context, where memory contains personal data, its design must be assessed within applicable data-protection obligations; this guide does not replace legal analysis or itself determine a lawful basis for processing.
Retention decision by information type
| Type | Purpose | Indicative retention | Invalidation |
|---|---|---|---|
| Conversational context | Complete a session | End of session or a defined short period | Closure, expiry, or applicable deletion request |
| Task state | Resume a pending workflow | Until the case is completed or escalated | Closure, cancellation, or ownership change |
| Persistent preference | Authorized personalization | A documented, reviewable period | Correction by the affected person or loss of purpose |
| Audit record | Investigation and accountability | According to documented policy | Restricted access; no silent alteration |
Designing for failure: stop, check, and recover
Tool, network, and external-dependency failures are normal. So are incomplete responses, timeouts, and ambiguous states: a call may have reached the receiving system even though the agent did not receive confirmation. A robust design does not assume that retrying is always safe. It must first know which operation was attempted, which outcome was confirmed, and which operations can be repeated without changing the final effect.
HTTP semantics distinguish idempotent operations, whose intended effect is the same after one or several identical requests, from operations that are not idempotent. This distinction helps design retries, but it does not remove the need to check business state. A technically idempotent request may still have unintended consequences if its parameters are wrong. For creation, payment, or sending operations, an idempotency key and a status lookup before retrying are generally more appropriate controls than blindly repeating the call.
The agent needs stop conditions. Set a limited number of attempts, timeouts, a call budget, and escalation criteria. Once a threshold is exceeded, the case should enter a review queue with the minimum necessary context: planned action, correlation identifier, received responses, and steps already performed. Retrying indefinitely can amplify an external incident or create duplicate actions.
Every meaningful action should have a checkpoint before its irreversible effect and, where feasible, a compensation plan. Compensation does not always equal a perfect reversal: refunding an amount does not erase a message already sent, and restoring a record does not remove a possible disclosure. Recovery documentation should state those limitations explicitly.
Recovery process for an uncertain outcome
- 01Assign a correlation identifier and record the intent before calling the tool.
- 02Apply a timeout and classify the error: validated rejection, transient failure, or unknown outcome.
- 03For an unknown outcome, query the state in the receiving system using the available identifier.
- 04Retry only if the operation and tool policy permit it; use an idempotency key where available.
- 05If the state cannot be confirmed, stop further related actions and escalate for human review.
- 06Record the resolution, any compensation applied where relevant, and the identified cause.
Observability and auditing without collecting unnecessary data
Observability makes it possible to reconstruct what happened; it does not require retaining every exchanged datum indefinitely. For a meaningful action, the record should associate the original request, applied policy, selected tools, authorized parameters or a protected representation of them, execution identity, outcome, retries, and human intervention. Correlation identifiers make it possible to follow a case across components without copying the full content into every log.
Logs must be protected as a sensitive asset. If they include prompts, responses, documents, or parameters, they may contain personal information, secrets, or untrusted instructions. They therefore need access controls, defined retention, environment separation, and mechanisms that prevent silent modification. It is also useful to distinguish operational logging for detecting failures from audit logging intended to investigate a decision; they may require different levels of detail and access.
A good reconstruction separates facts from interpretation. It should be possible to identify which data a tool returned, which rule blocked or allowed an operation, and which recommendation the model made. It is not reasonable to promise a complete explanation of all model behavior, but it is possible to record the chain of programmatic decisions and operational artifacts that determined whether an action was executed.
Minimization is not only a privacy obligation; it improves the security and usefulness of records. Excessive logging makes relevant signals harder to find and expands the data set exposed by improper access.
Evaluate before deployment: task, permissions, and recovery
Evaluation must reproduce the decision environment, not merely measure the quality of a textual response. A minimum plan combines representative tasks, edge cases, tool failures, ambiguous requests, attempts to inject instructions through external content, and authorization checks. Success criteria should include both correctly completing a permitted task and rejecting a prohibited action, stopping under uncertainty, and escalating when appropriate.
Benchmarks are useful for comparing particular capabilities under defined conditions. [Terminal-Bench](/benchmarks/terminal-bench) evaluates task completion in isolated terminal environments, while [OSWorld](/benchmarks/osworld) includes tasks across web and desktop applications in real evaluation environments. These measurements can provide signals about an agent’s performance on those tasks, but benchmarks do not replace testing in your own environment and do not certify that an identity has correct permissions, that organizational data is protected, or that the agent recovers safely in a specific integration.
Internal testing needs test accounts, synthetic or appropriately controlled data, dependency simulation, and recovery cases. It should also verify that an authorization denial does not trigger more permissive alternative paths. A risk-management framework can help assign owners, document decisions, and review controls throughout the lifecycle, but it does not replace the technical specification of each permission.
Test evidence, incidents, and tool changes should feed periodic reviews of the autonomy matrix. Capability results alone are not a release decision: the team must also establish that boundaries remain enforced when a tool fails, input is adversarial, or the agent encounters uncertainty.
Minimum cases for an evaluation suite
| Test | What is observed | Expected result |
|---|---|---|
| Normal authorized task | Accuracy and traceability | Completes the task within scope |
| Instruction embedded in a document | Resistance to instruction manipulation | Treats text as data and does not expand permissions |
| Parameter outside policy | Enforcement of limits | The tool rejects it or requests approval |
| Dependency timeout | Recovery | Checks status, limits retries, and escalates if needed |
| Simulated irreversible action | Human oversight | Generates a proposal and waits for explicit approval |
| Outdated memory | Invalidation | Prioritizes the current source or marks uncertainty |
Final template: deciding the autonomy level
The autonomy decision should be reviewable as a policy, not left implicit in a prompt. For every tool and action, document the risk category, involved data, specific permission, identity, limits, approval requirement, recovery strategy, records, and owner. If any of these elements is undefined, the assigned autonomy is probably premature.
A reasonable starting point uses four levels. Read-only permits queries against authorized resources. Proposal with approval permits investigation and preparation of an action, but reserves execution for a person. Limited execution enables bounded, reversible operations under deterministic rules. Autonomous execution is reserved for low-impact actions with limited scope, known reversal or compensation, available oversight, and sufficient test evidence.
The matrix should not be permanent. An incident, provider change, new tool, expansion of accessible data, or change in business policy justifies reviewing it. Likewise, an agent can begin in proposal mode and gain autonomy only after demonstrating reliable behavior within a measured scope. Reducing autonomy after an anomaly is a control response, not a project failure.
As a next editorial step, the conclusion can link to [deciding whether to automate a workflow](/choose/automatizar-flujos-con-agentes). The decisive question is not whether an agent can perform a task, but whether the organization can bound, observe, and recover its effects at an acceptable level of risk.
Autonomy matrix template
| Level | Can do | Cannot do | Requirement to advance |
|---|---|---|---|
| Read-only | Query assigned resources | Modify, send, or delete | Access logging and data filters |
| Proposal with approval | Prepare action and evidence | Execute without confirmation | An informed approval interface |
| Limited execution | Reversible operations within thresholds | Exceed scope, amount, or volume | Server-side validation, limits, and tested recovery |
| Autonomous execution | Predefined, low-impact actions | New, ambiguous, or irreversible actions | Monitoring, auditing, revocation, and periodic review |
Scope and uncertainties
This guide presents a technical and operational framework based on the supplied institutional sources, technical documentation, and agent-building guidance. Its architecture recommendations do not by themselves guarantee the safety of a use case or replace testing, threat analysis, privacy review, sector-specific controls, or legal advice.
The application of the General Data Protection Regulation depends on the circumstances of the processing, the roles of the parties, and other applicable requirements. In particular, this guide does not determine whether a specific case falls within rules on individual decisions based solely on automated processing or which additional safeguards are required. It also does not cover national, employment, financial, health, or contractual obligations that may be relevant.
Tool, model, API, and benchmark characteristics change quickly. Before deployment, the team should confirm the evaluated version, integration behavior, effective permissions, and execution conditions. Evidence from a test environment should not automatically be extrapolated to production.
Open questions
- Specific thresholds for amount, volume, retention, and escalation must be defined for each organization and use case; the sources do not establish universal values.
- Whether an action is reversible depends on the business process and its external effects, not only on a technical undo operation.
- This guide does not resolve the specific legal applicability of the GDPR or additional sectoral and jurisdictional rules.
- Benchmark outcomes and model capabilities alone cannot establish the safety of a production integration.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction