Ilustración editorial para Humanos en el circuito para agentes de IA: cómo diseñar aprobaciones que reduzcan riesgo sin bloquear el trabajo
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

The problem: “a human approves” does not define a sufficient control

An agent that can query internal systems, modify records, send communications, or initiate transactions does not cease to be risky merely because a person appears somewhere in the workflow. Oversight works only when human intervention is meaningful: the person must be able to understand the proposed decision, prevent it, correct it, or stop the process before it creates an effect beyond the accepted level of autonomy.

The phrase “requires human approval” often conceals decisive design questions. It does not say what is being approved—a goal, a plan, a specific action, or a batch. Nor does it identify the person authorized to approve, the information available to them, the maximum response time, or the default behavior if they remain silent. Without those definitions, approval can become a formality or, at the other extreme, a queue that prevents the workflow from delivering value.

It is useful to treat approval as a decision control, not as an interface. The control must connect an action class to a risk level, an accountable owner, a minimum body of evidence, an execution rule, and an auditable record. Technical permissions are still necessary: an approval should not expand the agent’s privileges or replace authorization by the destination system.

This approach complements the broader agent design covered in the learning section and should be coordinated with tool evaluation and discovery decisions. Its scope here is more specific, however: deciding where a person intervenes in an agent action and how to demonstrate afterward that the intervention was effective.

02

Three modes that must not be confused: prior approval, post-execution review, and an emergency stop

Prior approval suspends an action before it creates an external effect. It is appropriate when the potential impact is high, reversal is limited, sensitive data is involved, or there is not yet enough evidence that the agent acts reliably in that situation. Its cost is waiting time and reviewer workload, so it should not be imposed indiscriminately.

Post-execution review allows an action to run within predetermined limits and examines a sample, an alert, or the full set of results afterward. It is suitable when harm is bounded and reversible, a tested correction mechanism exists, and the organization can quickly detect unwanted effects. It does not mean the absence of control: it requires complete records, alert thresholds, accountable follow-up owners, and a real ability to reverse.

An emergency stop is a separate mechanism. It should make it possible to pause a specific execution, disable an integration, or temporarily withdraw autonomy for an action class. It is necessary even in workflows with prior approval, because systemic incidents, signs of manipulation, or contextual changes may make it inappropriate to continue. Agentic risk-management guidance recommends controls to guide, correct, and interrupt autonomous behavior, especially for high-impact actions or ambiguous inputs.

There is also a clarification intervention: the workflow pauses to request missing information, rather than to authorize an already well-specified action. Keeping it separate from approval prevents an informational reply from being mistakenly interpreted as consent to execute. Some workflow platforms implement human steps that can suspend a pending execution and request review or additional information; the technical pattern does not itself determine the risk policy.

Which mode fits each need

ModeQuestion it answersTimingMinimum condition
Prior approvalShould this action happen?Before the external effectAction and parameters are locked
Post-execution reviewWas execution within limits correct?After executionReversal and detection are available
ClarificationWhat information is missing to continue?Before planning or executionThe response does not authorize action on its own
Emergency stopShould the workflow or integration stop?At any timePause authority and procedure are defined
03

Build a decision matrix: impact, reversibility, sensitivity, scope, and empirical confidence

There is no universal list of actions that always require approval. Classification must start with the specific action and its context. In one organization, updating an internal label may be harmless; in another, the same change may trigger a service exclusion or alter a regulated record. The matrix should therefore document the effect that the action creates in the destination system, not just the name of the tool.

Assess at least five dimensions. Impact is the possible magnitude of harm to people, customers, operations, finances, or compliance. Reversibility measures whether the effect can be undone fully, safely, and at reasonable cost. Sensitivity covers the data that is queried, disclosed, or transformed. Scope considers volume, recipients, systems, and duration. Finally, empirical confidence is not an impression of the model: it is evidence from comparable testing and operations that the agent correctly identifies the case, proposes valid parameters, and does not omit relevant conditions.

A score can help organize decisions, but it should not automate a conclusion without veto rules. For example, an irreversible transaction or access to especially sensitive data may require approval even when the action concerns one item and the agent has performed well in tests. Conversely, a low-impact action may require prior review when it occurs in an anomalous context, the agent uses a new tool, or validation signals are unavailable.

The matrix should produce one of four policies: automatic execution within limits; prior approval by an accountable person; dual approval by distinct roles; or execution prohibited for the agent. The last category does not necessarily mean that the task is prohibited for the organization; it means that it requires a human procedure or a different integration.

Initial autonomy-policy matrix

Predominant signalsSuggested policyExample limitAdditional control
Low impact, reversible, limited scope, and stable evidenceAutomatic executionUpdate an unpublished internal draftLogging and post-execution sampling
Medium impact or contextual uncertaintyPrior approvalModify a single operational recordPreview and response deadline
High impact, sensitive data, or broad scopeDual approvalSend a large-scale external communicationSeparation of duties and enhanced logging
Irreversible, prohibited, or lacking reliable reversalDo not execute through the agentTransfer funds or delete evidenceRoute to a human process
04

Design the approval request so it can be reviewed

A reviewer should not have to reconstruct the agent’s full reasoning or rely on a persuasive explanation to make a decision. The request must present the facts and action boundaries in verifiable form. Its purpose is to make errors in the recipient, scope, data, authorization, or expected consequence detectable.

Include the operational objective, the exact action intended, the destination tool or system, binding parameters, the data that will be queried or disclosed, and a preview of the effect. When possible, present the safer or less intrusive alternative, such as saving a draft instead of sending it, limiting the batch, or requesting clarification. Also state what will happen if the person rejects, modifies, or does not answer the request.

The interface should clearly distinguish approving the action as defined, requesting changes, and rejecting it. A free-text field can be useful for an exception, but it must not allow an ambiguous instruction to become broad authorization. If the reviewer changes an essential parameter, the agent must create a new proposal or execute a validated deterministic workflow; it must not freely reinterpret the modification.

A good request also declares the provenance of the context: which internal sources or user inputs support the proposal, which checks were performed, and which uncertainties remain open. Showing this information does not mean exposing unnecessary data to the reviewer. Visibility must respect data minimization and the access restrictions of the role itself.

05

Assign accountable owners, alternates, and escalations

The person reviewing must have real authority over the action’s effect. An operations owner may approve an inventory update within their remit, while access to personal data may require a data owner or security role. Assigning reviewers solely based on availability commonly produces mechanical approvals or rejections caused by lack of context.

Document a policy owner for every action class, the roles authorized to approve, and the conditions requiring separation of duties. Dual approval makes sense when one person should not fully control a decision: for example, one role understands the operational need while another validates compliance risk. It should not be the automatic response to every uncertainty, because duplicating reviews without distinct purposes can increase delay without increasing detection.

Define alternates with the same level of authority, but avoid automatically forwarding a request to many people. A queue with named owners, a deadline, and explicit escalation is more auditable. A request should expire if the conditions supporting it change, such as the validity of a data item, the status of a case, or the content of an integration.

AI risk-management frameworks recommend assigning roles, responsibilities, and authorities, documenting risk, and maintaining post-deployment monitoring. In a system with agents, that assignment must cover both the decision to allow an action and the ability to intervene during an incident.

Escalation process for a pending request

  1. 01Create the request with an action identifier, locked parameters, and an expiration date.
  2. 02Notify the domain owner and record the delivery time.
  3. 03If there is no response by the first threshold, alert the authorized alternate without expanding the action.
  4. 04If the request expires, apply the safe default policy: do not execute, preserve the context, and close or return the case to a queue.
  5. 05Escalate to the policy owner when the lack of response affects a critical service or reveals insufficient review capacity.
06

Handle silence, rejection, and disagreement without creating bypasses

Silence is not approval. The default rule should be not to execute when an action is awaiting authorization and the deadline expires, unless a documented policy establishes a safe alternative action. For example, an agent may save a draft, create a task, or request additional data, but it must not turn a lack of response into permission to send, change, or disclose.

A rejection must have defined semantics. It may close the case, return it to the agent with explicit parameters that it must honor, or route it to a person for manual execution. It is important to prevent the agent from retrying the same action with superficially different wording to obtain a new approval. Group equivalent retries, preserve the link to the prior rejection, and require an identifiable material difference.

Disagreements between approvers also need an outcome: precedence for a risk role, referral to a case owner, or a block until a person with higher authority makes a decision. The system must retain decisions and their operational rationale without assigning certainty to an explanation generated by the agent.

After approval, verify that the executed artifact matches the approved one. Compare the action identifier, plan version, parameters, recipients, tools, and result. If any of these elements changes, request a new approval or apply a previously authorized, tightly limited change path.

07

Patterns by action type and execution boundaries

For external communications, separate draft generation from sending. It may be reasonable to automate classification and preparation while content remains internal; sending requires a higher threshold when it affects customers, commits a contractual position, or contains sensitive data. Changes to recipients, language, attachments, or distribution lists should be treated as material changes.

For code or configuration changes, human review of content does not replace delivery-cycle controls. Approval must be tied to an identifiable version, available tests, the destination environment, and a rollback plan. An agent should not expand deployment from a test environment to production through approval obtained for limited validation.

For record updates, establish which fields the agent may modify, which sources are admissible, and when it must show the difference between current and proposed values. Bulk updates require a specific control over scope, selection criteria, and the undo mechanism. For data access, apply least privilege, authorization in the destination system, and execution in the authorized context; approval at the agent layer must not grant access that the source system denies.

Transactions with financial, legal, or physical effects generally require a restrictive policy. Classification depends on context and existing controls, but irreversibility, possible impact on third parties, and difficulty of remediation are strong signals to require dual approval or exclude the agent from execution. Security recommendations on excessive agency emphasize least privilege, authorization in the destination system, oversight of high-impact actions, and activity logging.

Control patterns by action type

Action typePrudent initial autonomyEvidence needed to approveChange that invalidates approval
External communicationAutomatic draft; send according to riskRecipient, text, attachments, and included dataRecipient, content, or list
Code or configuration changeProposal and tests; controlled deploymentVersion, environment, test results, and rollbackVersion, environment, or scope
Record updateLimited fields and volumeBefore and after values, source, and selection criterionField, batch, or source
Data accessOnly permissions already grantedPurpose, dataset, and durationData, purpose, or identity
Sensitive transactionEnhanced approval or human executionAmount, counterparty, conditions, and effectAny material parameter
08

Measure whether the control detects real errors and adjust autonomy using evidence

The existence of approvals does not demonstrate that the control works. Measure how many proposals are rejected or modified because of material errors, which error types are detected, how many approved actions must be reversed, and how long the queue takes. Segment these metrics by action class, integration, workflow version, and owner, because a global average can conceal a concentrated problem.

Approval rate alone is ambiguous. A very high rate may reflect correct, low-risk proposals, but it may also reflect fatigue, lack of context, or pressure to clear the queue. Investigate combined signals: approvals completed unusually quickly, repetitive comments, discrepancies found in post-execution review, reversal rate, and differences between reviewers. Quality samples and negative-case reviews help distinguish efficiency from improper automation.

Agent confidence must be updated using observed results, not a general performance claim. To move an action class from prior approval to post-execution oversight, define in advance the evidence required: a comparable case population, material errors below the internal threshold, proven reversal, post-execution detection within an acceptable time, and no relevant changes to the model, tools, or data. If those elements change, reassess the policy.

Record a trace that allows the case to be reconstructed: the original request, permitted context, plan, proposed action, agent and tool versions, approver, decision, time, executed parameters, destination-system response, and subsequent outcome. The record must be protected and accessible only to authorized roles; collecting more context than necessary can create an additional privacy or security risk.

Metrics and interpretation signals

MetricWhat it may revealInterpretation riskFollow-up action
Material errors detected before executionPreventive value of reviewA low volume may indicate poor detectionAudit samples and subsequent cases
Reversal or rollback rateFailures that passed the controlIt can vary with difficulty of reversalAnalyze by action and cause
Queue time and expirationsActual response capacityShortening the deadline does not solve missing ownersAdjust coverage and escalation
Very fast, repetitive approvalsPossible fatigue or automationThey may also correspond to simple casesReview displayed evidence and sample decisions
Actions outside policyBoundary or integration defectsUnderreporting may occurReconcile logs with destination systems
09

Implementation plan and verifiable launch checklist

Start with a narrowly scoped workflow, a single action class, and a destination system where the result can be observed and reversed. Inventory the workflow’s actions, not just its tools: querying, drafting, updating, sending, deleting, escalating, and executing can require different controls. For each one, describe the effect, technical permissions, parameter limits, and policy owner.

Before deployment, test normal cases, ambiguous inputs, malicious requests, tools that return errors, and parameter changes after approval. Verify in particular that the agent cannot execute an action different from the approved one, that an expired request does not execute, and that the emergency stop affects both pending and new executions.

Run a controlled phase with enough evidence to detect patterns, without claiming that a fixed number of cases is valid for every risk. Review rejections, reversals, wait times, unanswered requests, and incidents weekly. Expand autonomy only for a specific action class and only when results and reversal mechanisms justify the change.

The approval policy is a living document. It must be updated when tools are added, available data changes, the model is modified, a new destination system is integrated, or incidents occur. Governance must retain the version of the policy applicable to every execution so that the historical record can be interpreted correctly.

Launch checklist

  1. 01List every action with an external effect and describe its proven reversal or the absence of reversal.
  2. 02Classify impact, reversibility, sensitivity, scope, and empirical confidence; document veto rules.
  3. 03Assign a policy owner, approvers, alternates, and dual-control cases.
  4. 04Design the request with the action, parameters, affected data, preview, alternatives, and expiration behavior.
  5. 05Lock or version approved parameters and verify that changes invalidate approval.
  6. 06Test rejection, silence, disagreement, tool errors, and the emergency stop.
  7. 07Record the proposal, decision, identity or role, executed parameters, and subsequent outcome.
  8. 08Define metrics, review frequency, and explicit criteria to increase or reduce autonomy.

Open questions

  • There is no universal threshold for deciding when an action should move from prior approval to post-execution review; it must be defined according to context, reversal capacity, and operational evidence.
  • Approval deadlines and dual-control requirements depend on service criticality, role authority, and obligations applicable to each organization.
  • The proposed matrix is an operational starting point and does not replace legal, privacy, security, or destination-system-specific policy analysis.
10

Keep exploring

10

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction