Ilustración editorial para Despliegues graduales de IA: cómo lanzar cambios de modelo, prompt o herramienta sin convertir a los usuarios en el experimento
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

An AI change is not just changing a model

In a production AI application, a modification may affect the model identifier or version, the system prompt, included examples, generation parameters, the orchestration chain, the information-retrieval policy, the queried index, an external tool definition, the permissions granted to that tool, or retry logic. Even when the change appears local, it can alter the final response, cost, latency, security behavior, and the likelihood of executing an external action.

The unit to deploy and evaluate is not necessarily the model in isolation. In a generative application, behavior results from the combination of the model, instructions, retrieved context, tools, and application rules. For example, replacing a model in a retrieval-enabled assistant can change how it interprets instructions and uses retrieved documents. Changing a tool schema can make a previously valid response no longer executable even when the generated text appears correct.

For that reason, the goal of a gradual deployment is not to show that a new version works in a few cases. It is to reduce uncertainty before exposing a broad population to a regression. The team formulates a testable hypothesis, compares the variant against a frozen baseline, initially limits the blast radius, and makes predefined decisions using operational data and quality evaluation.

This guide addresses controlled change once an application is already operating. It does not replace the design of an evaluation set, initial observability instrumentation, or the strategic selection of a provider or model. It is useful alongside the learning, comparison, and discovery resources available through the internal routes learn.index, compare.index, and discover.index.

Typical blast radius by modified artifact

ArtifactEffects worth checkingCan a limited comparison be enough?
Model version or familyTask quality, format, safety, latency, cost, and tool useSometimes; it depends on functional equivalence and the permissions involved
Prompt or examplesInstruction adherence, data extraction, tone, format, and refusal behaviorYes, if the output contract is preserved and the action scope does not change
Index, corpus, or retrievalCoverage, freshness, internal citations, confidentiality, and incorrect contextOften requires new cases and segmentation by source
Tool definition or permissionsArguments, actions, duplicates, authorization, and reversibilityNot on its own when it can create external effects
Retries, timeouts, or fallbackLatency, cost, duplicated actions, and real success rateYes, with load testing and monitoring of side effects
02

Before moving traffic: hypothesis, baseline, and metric types

A canary does not fix a poorly framed decision. Before enabling it, describe the change concretely: which artifacts change, what remains fixed, which population could be affected, and which outcome is expected to improve. A useful hypothesis avoids statements such as “the model will be better.” Instead: “variant B will increase the validated resolution rate for billing queries without increasing incorrect tool calls or exceeding the latency budget.”

Then freeze a baseline. Record a time window, the included traffic, the exact configuration versions, and the observed indicators. If data from periods with very different demand, case mix, or tool availability are compared, attributing a result to the change will be uncertain. When possible, assign control and treatment simultaneously to reduce that difference; when that is not possible, state the limitation and interpret results cautiously.

Classify metrics into three groups. Promotion metrics determine whether traffic may grow: for example, validated success per task, compliance with a format, or accuracy reviewed in a sample. Monitoring metrics detect effects that should not worsen materially, such as queue latency, cost per completed task, abandonment, or fallback rate. Rollback metrics are safety, privacy, incorrect-action, or operational-degradation limits that require stopping without waiting for statistical analysis to finish.

Do not use a global average as the only criterion. An aggregate improvement can hide severe deterioration in long queries, less common languages, accounts with limited permissions, or tasks that invoke a tool. Break down metrics by usage segment, complexity, tool path, and outcome. The need for representative participation and caution when extrapolating evaluation results to the real context are particularly important in generative AI.

Minimum preparation before exposing the variant

  1. 01Define the change package and assign it an immutable identifier.
  2. 02Write the hypothesis, target population, excluded segments, and decision owner.
  3. 03Freeze the baseline using the same metric definitions that will be used during the canary.
  4. 04Set promotion, pause, and rollback thresholds, together with the action associated with each one.
  5. 05Verify that the flag can return to the prior package without manually editing configuration.
  6. 06Prepare event logging, output samples, and the human-review procedure with access controls.
03

Classify the change before choosing the mechanism

Classification does not certify that a change is safe; it helps determine how much additional evidence is needed. A minor modification preserves the model, input and output contract, tools, permissions, and context source, and changes a bounded aspect, such as a formatting instruction. It can proceed with offline replay, a review sample, and a small canary if it does not touch sensitive actions.

A comparable change replaces one component but preserves an equivalent task, interface, permission set, and definition of success. An alternative model for document summarization, using the same prompt, retrieval, and structured output, may fall into this category. Even so, the team must measure cost, latency, variation, and formatting failures; equivalence is a hypothesis to test, not a property declared by the provider.

A change requiring a new evaluation alters what the system can do, the information it can access, or the potential harm. It includes adding a tool that creates tickets, expanding permissions, switching to a corpus with different data, allowing irreversible actions, changing the served population, or introducing a retrieval policy that transforms sensitive material. In these cases, randomized traffic may be insufficient or inappropriate. Specific approval, controlled testing without external effects, and review of applicable requirements may be needed.

Documenting artifacts is part of classification. For every package, retain the model identifier, prompt and templates, inference parameters, code version, orchestration chain, retrieval configuration and index or corpus snapshot, definition and version of each tool, permissions, input and output schemas, retry policy, and feature-flag rules. Without this information, it is not reasonably possible to reproduce a response or confirm what was rolled back.

04

Choose replay, shadow, canary, A/B, or segment-based deployment

Offline replay reruns historical requests, subject to appropriate privacy restrictions, against the candidate variant. It is inexpensive and reproducible for comparing format, retrieval, estimated cost, and routing decisions. It does not perfectly reproduce real interaction, state changes, or the behavior of external tools. Use it to rule out obvious failures, not to declare that production impact has been proven.

In shadow mode, the variant receives a copy of real requests, but its result is neither shown to the user nor allowed to execute external effects. It is suitable for observing latency, cost, output stability, tool selection, and divergence from the active system. To remain safe, tool calls must be simulated, redirected to an isolated environment, or blocked. Shadow mode stops being representative if the user would have provided clarification after seeing the response, if the variant needs context it does not receive in parallel, or if a tool depends on mutable state created during the conversation.

A canary sends a limited fraction of traffic to the new version and increases it in stages if checks are met. It is useful when the response can be presented to the user with bounded risk and there is a fast rollback path. The initial size should not be set by habit: it must allow relevant operational signals to be detected within an acceptable exposure limit. Its duration must cover representative patterns, including high-load hours, without keeping the experiment open longer than necessary in the presence of an adverse signal.

An A/B test can estimate differences between variants if assignment is stable, populations are comparable, and the metric is well defined. It is not synonymous with a canary: a canary prioritizes limiting risk during delivery, whereas an A/B test aims to attribute a difference to a variant. They can be combined, but there is no need to force randomization when ethical, regulatory, or safety restrictions apply. Segment-based deployment, meanwhile, makes it possible to start with lower-criticality cases or internal users, but it can introduce bias: a good result there does not guarantee the same result for the rest of the population.

05

Design a canary that produces evidence and limits harm

Before starting, define who can initiate, pause, promote, and roll back. Establish an accountable on-call function for each phase and an escalation channel. Every traffic increase should have an observation window, automated checks, and an explicit review of relevant signals. Do not promote automatically just because there are no alerts: the absence of an alert may reflect a poorly instrumented metric, insufficient volume, or lack of coverage for an important segment.

Assignment should be stable for the same unit, such as an account, organization, or conversation, unless there is a documented reason to use another unit. Switching variants in the middle of a conversation complicates both the experience and diagnosis. Also avoid initially including vulnerable populations, high-impact processes, or accounts whose configuration makes a clean rollback impossible. That exclusion reduces exposure, but it also reduces representativeness; the plan should state when and under which controls those segments will be evaluated.

Define spending and capacity limits. A variant may produce longer responses, make more calls, or trigger retries that increase cost even though apparent quality improves. Monitor both cost per request and cost per completed task, because reducing unit cost at the expense of more abandonment is not necessarily an improvement. Also measure queues, dependency errors, high-percentile latency, and saturation of retrieval services or tools.

In non-deterministic systems, one execution is not enough to characterize a critical case. Repeat a subset of inputs with the same configuration and measure the spread of relevant results: structural validity, the decision to use a tool, policy compliance, or human score. That spread should not be interpreted as an exact probability when sampling is small or conditions change; it is a signal to broaden testing and establish safeguards.

Decision rules that must exist before launch

SignalInitial actionCondition to continue
Security, privacy, or unauthorized external-action failureStop the increase and roll back if impact is not containedInvestigation, correction, and new validation of the package
Sustained degradation in a promotion metricPause the phaseReviewed evidence that the difference is within the agreed threshold, or a correction is applied
Increase in cost or latency without critical harmDo not increase traffic; analyze configuration and loadCost and latency are within operational limits without degrading task success
Invalid output in critical casesRemove the variant from that flow or roll backSchema, prompt, model, or validation corrected and retested
Consistent improvement with no alerts or pending exclusionsPromote to the next phaseObservation window completed and the accountable owner authorizes progress
06

Promote, pause, roll back, or retire: explicit decisions

Promoting does not mean declaring that the change is universally better. It means that, for the observed segment and phase, the evidence meets the defined thresholds and there are no signals advising a stop. Record which data were reviewed, which segments remain missing, and which uncertainties are accepted. This prevents a gradual promotion from becoming an ownerless expansion through inertia.

Pause when there is an ambiguous signal: insufficient volume, an unexpected traffic distribution, unavailability of a dependency, or a difference that requires human review. Pausing preserves the exposure limit while the cause is clarified. It must not be used to ignore a high-impact alert; when a rollback threshold is crossed, the action is to roll back or disable the affected capability.

An effective rollback is executed through a known, tested reference to the prior package, not by rebuilding prompts or configurations under pressure. The reversal must cover every linked component: model, prompt, parameters, retrieval, tools, schemas, permissions, retry rules, and flag. If a data migration or external action cannot be undone, that irreversibility must be part of the pre-deployment analysis and approval controls.

After a rollback, retain enough evidence to investigate: assigned version, time, segment, request with appropriate minimization or pseudonymization, permitted retrieval context, output, tool calls, validator result, latency, cost, and the decision made. Access to these records must comply with applicable security and retention controls. Recording more data than necessary can create privacy risks; recording less can prevent diagnosis.

Permanently retiring a variant is also a valid decision. If the change does not demonstrate operational benefit, persistently increases risk, or requires disproportionate controls, document the result and close the experiment. Useful learning includes knowing which hypotheses were not supported.

Launch-plan template and minimum decision dashboard

  1. 01Package: immutable identifiers for model, prompt, parameters, retrieval, tools, permissions, and code.
  2. 02Hypothesis: expected improvement, promotion metric, population, and conditions held constant.
  3. 03Risk: external effects, irreversibility, excluded segments, spending limits, and required approvals.
  4. 04Phases: replay, shadow, initial percentage, increments, duration, and owner for each gate.
  5. 05Dashboard: task success, validation errors, security events, tool use, latency, cost, fallbacks, and breakdown by segment.
  6. 06Decision: promotion, pause, and rollback thresholds; authorized person; recorded time and justification.
  7. 07Post-investigation: permitted samples, retention, findings, corrections, and final decision to expand or retire.
07

Limits and questions that must remain open

No protocol eliminates the uncertainty inherent in a generative application. Canary results may not generalize to periods of greater load, new query types, different languages, or subsequent changes in dependencies. Benchmarks and internal tests do not replace observation in the operational context either. That is why monitoring and rollback capability should remain in place after reaching one hundred percent of traffic.

Statistical significance, where it applies, does not replace operational judgment. A small but statistically detectable change may have no practical importance; a rare security event may require rollback even without enough volume for conclusive calculations. Thresholds should reflect the severity of harm, reversibility, and the context of use, not merely a numerical difference.

There is also uncertainty about the behavior of third-party managed services and models: updates, capacity limits, latency changes, or output variation can affect the result. Versioning what the team controls and recording the versions or identifiers exposed by the provider improves traceability, but it does not make the environment fully deterministic. The plan should identify external dependencies and how their changes will be detected.

The final criterion is easy to state and demanding to apply: expand only when the change demonstrates sufficient value within agreed risk limits; pause when the evidence does not allow the result to be interpreted; roll back when a harm limit is crossed; and retain the prior package until the new behavior is sufficiently understood. In this way, the user stops being the primary mechanism for discovering failures and is instead protected by a deliberate delivery process.

Open questions

  • The initial size and duration of a canary have no universal value: they depend on volume, severity of harm, task variability, and the capacity to intervene.
  • Shadow mode may stop representing real usage when user interaction is missing, external state changes, or tools with real effects are blocked.
  • Limited-traffic results may not generalize to excluded segments, demand peaks, new languages, or changes in third-party dependencies.
  • Repeating executions helps observe variation, but does not guarantee a conclusive statistical estimate when the input set is small or unrepresentative.
  • Regulatory, privacy, and approval obligations vary by sector, jurisdiction, and use case; they must be reviewed for every deployment.
08

Keep exploring

08

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction