A Demo That Responds Does Not Prove Compatibility
In an AI application, a change may produce no transport error and still break the product. The model may continue to return text that is useful to a person while omitting a field consumed by a downstream service, selecting a disallowed tool, changing the practical meaning of a label, or grounding a conclusion in a retrieved passage that does not support it. Therefore, checking that it “still responds” is not the same as checking that it preserves the expected operational behavior.
A contract is a verifiable specification of the properties that must be preserved at a specific boundary in the workflow. It is not a promise that the model will always write the same sentence, nor an attempt to eliminate generative variability. It defines which inputs are accepted, what shape the output must take, which business invariants cannot be violated, which actions are authorized, and what evidence is required before making a claim or carrying out an operation.
The unit of analysis should be the concrete incompatible change. It may be a model replacement, an SDK update, a system-prompt modification, a schema change, a new tool definition, or an updated provider policy. The goal is to answer before promoting the change: has the contract been preserved? Is the degradation within an accepted limit? Does the workflow need to be adapted? Or must the change be blocked?
Contracts complement aggregate quality evaluations, but they solve a different problem. An evaluation may show that average usefulness remains high even though one out of one hundred cases issues a cancellation order without confirmation. That isolated case is a critical incompatibility if the action has external effects. Likewise, valid JSON does not prove that an argument is safe, that a citation corresponds to its source, or that a classification retains its meaning.
Inventory of Boundaries and Dependencies
Before writing tests, map the complete path of a user case. A boundary exists wherever one component hands over a representation that another component interprets: the request sent to the provider, the model identifier, the model response, the SDK adapter, the context retriever, a tool call, the destination system, and the external action. Each boundary may have a different contract and a different owner.
The inventory must identify effective versions and configurations, not merely generic names. Record the requested model or snapshot, SDK and API version where applicable, system prompt, generation parameters, output schema, tool list, permission definition, retrieval-index version, and internal adapters. Without this information, a later failure can scarcely be attributed to a specific cause.
Not every boundary requires the same type of assertion. The model request requires checking accepted parameters and normalized values. A structured output requires validation of its schema and mandatory fields. Retrieval requires checks for provenance, freshness, and sufficient evidence. Tools require authorization and effect simulation in addition to syntax. The destination system requires idempotency, transaction control, or compensation according to the risk.
Provider documentation indicates that model life cycles and interfaces can change. In particular, model retirements and certain parameter changes can turn previously valid requests into errors. This makes it necessary to test a migration as a dependency change, even if the business code has not changed.
Boundaries and minimum checks
| Boundary | Minimum contract | Failure it reveals |
|---|---|---|
| Request and SDK | Accepted model, parameters, and serialization | Retired parameter or modified format |
| Structured output | Schema, types, fields, and permitted values | Missing field or unexpected enumeration |
| Retrieval | Document, date, authority, and sufficient evidence | Answer not supported by the context |
| Tool | Permitted tool, arguments, and authorization | Action with the wrong scope or data |
| External destination | Preconditions, idempotency, and logging | Duplicate or irreversible effect |
What Should Become a Contract
Start with deterministic elements. Good candidates include types, required fields, numeric limits, enumerations, identifiers, the presence of a source, date formats, permitted tools, and authorization rules. Business invariants also belong here: a refund cannot exceed the amount paid, an agent cannot modify another customer's data, and an operation that requires human approval cannot run without that state.
Then add operational constraints. Define a latency and cost budget per case, with a specified measurement method: for example, a percentile over a controlled sample rather than an isolated impression. Establish maximum retries, tool calls per execution, retrieved documents, and context size. An increase may be technically compatible yet unacceptable for the product; the contract must separate those two dimensions.
Labels deserve explicit semantic treatment. If an output contains `high_risk`, the contract must explain which facts justify it and which consequences it triggers. Literal agreement on the label is insufficient if the assignment criterion has changed. Use boundary cases with human annotation and assertions about the observable conditions that should lead to each class.
For evidence-backed answers, the contract must distinguish between having a citation and being supported. At a minimum, check that the retrieved source is admissible for the domain, that its date satisfies the freshness policy, that the passage contains sufficient evidence for the claim, and that the workflow declares insufficiency when it lacks a basis. Attribution must not become a decoration generated at the end of the process.
Classify the Change Before Debating Its Results
Classify every proposed modification into four groups. A compatible change preserves all applicable contracts. A change compatible with acceptable degradation misses a non-critical target within an approved threshold, such as a limited latency variation. An incompatible change violates a mandatory property, such as an action permission or a required field. An unknown change is one for which cases, fixtures, telemetry, or a sufficiently precise definition are missing to reach a conclusion.
The classification must not depend on who proposes the change or on whether a demonstration looks convincing. It must be tied to previously published promotion rules. If the team discovers that a rule no longer reflects the product's needs, it may change the contract, but that decision must be explicit, reviewed, and versioned; it must not be implicitly accepted through a failed test.
Pinning a specific model version reduces one source of variation and makes results easier to reproduce. Aliases or models subject to updates can modify behavior without a change in client code. OpenAI documentation recommends pinning model versions and running evaluations because snapshots can vary in prompting behavior. Accordingly, a contract should record both the requested identifier and the update policy accepted by the team.
Promotion decision
| Result | Example | Decision |
|---|---|---|
| Compatible | Schema, permissions, and thresholds are preserved | Promote with the test record |
| Acceptable degradation | Latency increases within the approved budget | Promote and monitor the indicator |
| Incompatible | The tool receives an argument that violates a business rule | Block and correct or adapt |
| Unknown | There is no fixture for a new external action | Do not promote until evidence is available |
Designing a Minimum Test Suite That Is Diagnostic
A useful test suite does not need to represent every possible human conversation. It must contain fixed cases covering critical paths, edge cases, and historical counterexamples. Each case should declare its input, initial state, workflow configuration, expected outcome, severity, and assertions. Keep test data free of sensitive information and ensure it can be run repeatedly.
Use fixtures for tools and external dependencies. A fixture should return controlled states, record calls, and prevent real effects. This lets you verify that the model selected the right tool, that arguments were interpreted by the adapter, and that no prohibited alternative was attempted. A test environment that calls production is not a fixture: it mixes compatibility testing with operational risk.
Snapshots are appropriate for deliberately stable artifacts, such as a normalized request, a tool schema, or an ordered list of retrieved identifiers. They are fragile for complete prose generated by a model. For natural language, prefer bounded semantic assertions: required facts are present, prohibited claims are absent, statements correspond to evidence, and the system abstains when data is insufficient.
Some measurements are not deterministic. Success rate, latency distribution, and classification frequency may require several runs, a fixed sample, and a predefined interval or tolerance. Do not turn a small statistical difference into a critical regression, and do not allow statistical uncertainty to hide a deterministic safety violation.
Process for building the initial suite
- 01List the actions and decisions whose errors have material impact.
- 02Write one verifiable property for every critical assumption, with a severity level and owner.
- 03Create nominal cases, edge cases, and cases that previously produced incidents.
- 04Replace tools and destinations with observable fixtures that have no effects.
- 05Separate deterministic validations from metrics with statistical tolerance.
- 06Run the suite against the current reference before evaluating the proposed change.
Structured Outputs and Tool Calls: Schema Is Not Authorization
A structured output must be validated twice: first against its representation and then against its meaning. The first validation checks JSON, types, required fields, ranges, and permitted values. The second checks relationships between fields and external state. For example, that a start date precedes an end date, that an amount belongs to the specified order, and that a reason code is consistent with the case.
Provider tool interfaces may describe parameters through JSON Schema and offer strict modes for schema adherence. This is a meaningful aid for reducing malformed arguments, but it does not replace application validation. An argument may be correctly typed while referring to the wrong account, an operation outside policy, or an action that needs approval. The executor must apply authorization, preconditions, and limits before producing effects.
The contract must also establish ordering. In a workflow that checks eligibility and then issues a refund, do not accept an inverted sequence merely because both calls are individually valid. Record the tools permitted at each stage, the maximum number of invocations, normalized arguments, fixture response, and the absence of unauthorized calls. This makes it possible to detect changes where the model appears to solve the task but takes a dangerous operational shortcut.
If the provider changes the model, tool definition, or SDK adapter, run the same fixtures. A valid-JSON test would detect a malformed object; this suite can detect that a different tool was selected, that the prior check was skipped, or that the system attempted to repeat an already confirmed action.
Contracts for Retrieval and Source-Backed Answers
In a RAG workflow, the contract begins before drafting. Establish which collections the case may query, which minimum metadata every passage must return, and how freshness is resolved. If an answer depends on a current policy, an old document may be technically retrievable but inadmissible as support. The test must inspect provenance, not only the final text.
Define minimum evidence for each type of claim. A normative conclusion may require an explicit passage from an authorized source; a synthesis may require several consistent passages; a number may require an exact match with the document. If results do not satisfy that condition, the correct behavior may be to request more context, state uncertainty, or decline to make the claim. That abstention is a contractual output, not a default experience failure.
Test contradictions and insufficient context. Include fixtures with outdated documents, lower-authority sources, passages that mention similar terms, and sets containing conflicting information. The contract should state whether the workflow prioritizes one source, exposes the conflict, or escalates for review. It is not reasonable to say that a citation is correct merely because it shares words with the answer.
For every execution, retain the candidate-document set, the selected documents, their identifiers and relevant metadata, the index version, the transformed query, and the final result. This telemetry makes it possible to distinguish whether a violation arose in retrieval, in model interpretation, or in evidence representation.
CI/CD Integration, Decision-Making, and Rollback
Run the test suite whenever any registered artifact changes: model version, SDK, system prompt, parameters, schema, tool definition, retriever, index, or policy. The change should generate a comparable manifest including versions, hashes or internal identifiers, per-case results, duration, measured usage when available, and tool and retrieval traces. Do not allow an implicit update to escape change control.
Organize CI gates by severity. Critical assertions, such as authorization, disallowed effects, customer isolation, or mandatory evidence, must block. High-severity assertions normally block until an approved adaptation exists. Quality or performance metrics with tolerance may require review. An unknown result must not automatically become compatible because there is no signal.
When a contract fails, first locate the boundary. Compare the normalized request, the model artifact, the raw response, SDK adaptation, retrieved documents, and tool log. Then choose whether to correct the integration, adapt the contract because a legitimate need changed, version the workflow to preserve both behaviors, or reject the change. Document why the decision is valid and who approved it.
Rollback must be designed before promotion. Preserve the prior configuration that can restore a compatible model, prompt, schema, tools, and adapters. If the provider retires a version, an exact rollback may not exist; in that case, the alternative is a versioned workflow with a tested adaptation. Provider-published retirement policies are an additional reason to plan migrations before the deadline rather than after detecting an incident.
Triage of a violation
- 01Stop promotion if a blocking assertion fails.
- 02Identify the case, contract, version, and boundary that differ.
- 03Reproduce using the same fixture and recorded configuration.
- 04Determine whether it is a regression, a test defect, or a legitimate requirement change.
- 05Apply a correction, versioned adaptation, or rollback.
- 06Record the decision, residual risk, and review date.
Contract Manifest Template for Each Workflow
A concise manifest turns intent into a reviewable artifact. It should live alongside the workflow and change through the same review process as the code. It does not need to contain secrets or every test conversation; it should point to internal fixture identifiers and precisely define the properties it governs.
Include the workflow's name and purpose; technical and product owner; contract version; model, SDK, and configuration identifiers; prompt or reference to its version; output schema; permitted tools and permissions; retrieval dependencies; case list; performance and cost thresholds; severity of every rule; promotion policy; required telemetry; rollback strategy; and review date. If a rule has no owner or failure criterion, it is not yet an operational contract.
The template does not eliminate technical judgment. The available sources document versioning, retirement, schema, and strict tool-use mechanisms, but they cannot decide which evidence is sufficient for your domain or which action requires human approval. Those decisions belong to the responsible team and must be expressed as testable policies. The manifest's advantage is that it forces them to become visible before an update contradicts them.
Minimum manifest fields
| Field | Expected content |
|---|---|
| Identity | Name, version, owners, and review date |
| Dependencies | Model, SDK, prompt, schema, tools, and index |
| Rules | Invariants, permissions, evidence, cost, and latency |
| Tests | Cases, fixtures, severity, and tolerances |
| Operations | Promotion gate, telemetry, and rollback |
Open questions
- The supplied sources describe behaviors and mechanisms of specific APIs, but they do not establish a universal policy for severity, latency, cost, or sufficient evidence across all domains.
- Model availability, names, and retirement dates may change; the team must check current documentation before a migration.
- Strict adherence to a schema reduces formatting failures, but does not by itself guarantee factual correctness, semantic authorization, or the complete absence of unwanted effects.
- Tests involving generative models may retain residual variability even with pinned configurations; statistical thresholds should be calibrated with data from the specific workflow.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction