Ilustración editorial para Evaluación de conversación GPT‑Live: cómo medir turnos, interrupciones y resolución de tareas sin confundir una charla fluida con un agente fiable
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

A smooth conversation does not prove that the agent is reliable

Evaluating a full-duplex voice agent requires separating outcomes that are often blended together in a demonstration. A response that arrives at a good pace, uses plausible pauses, and tolerates some overlap can create the impression of a natural conversation. That impression does not, however, prove that the system correctly recognized an amount, understood a self-correction, retained the user’s latest instruction, or completed an external task without unintended effects.

The distinction is especially important when evaluating GPT‑Live‑1. The provider documentation describes a voice model that can participate in full-duplex interaction and delegate work to other models or tools. The observed outcome therefore does not depend on the conversational layer alone: turn detection, transport, transcripts when they are used, the delegated backend, tool definitions, permissions, and the state of external systems also matter. A final failure should not automatically be attributed to the voice model.

The evaluation should answer an operational question: given a specific spoken input and a specific system state, did the service understand what mattered, manage the turn appropriately, apply the correct policy, and leave the external system in a valid state? Naturalness may be a desirable outcome, but it does not replace that verification.

Full-duplex dialogue benchmarks support explicitly measuring turn-taking phenomena such as pauses, brief listening signals, interruptions, and overlap. Task-oriented benchmarks, by contrast, use completion or success criteria. These are complementary perspectives: no single rate rigorously summarizes both kinds of behavior.

02

Define the test unit before measuring

The test unit should not simply be a call or a transcript. It should be defined as a reproducible scenario: user profile, initial goal, input script or audio, temporal events, initial backend state, permitted tools, confirmation policy, cancellation policy, and success condition. This design makes it possible to repeat an execution and determine what changed when the result varies.

Each run should record the exact identifier of GPT‑Live‑1 or the variant used, the date, the region or environment when relevant, the transport channel, and the turn-detection configuration. The Realtime reference includes detection configurations based on server voice activity and semantic criteria, along with parameters that affect the sensitivity or speed of the decision. Comparing percentages without fixing these options blends operationally different systems.

The delegated chain must also be described. If the voice layer sends a request to another model, that component may determine reasoning, a tool call, or the wording of an argument. Record the backend version and configuration, available tools, argument schemas, timeouts, retries, permissions, and state sources. The GPT‑Live system card warns that the capabilities and safeguards of delegated work depend on the model or backend to which it is delegated.

For cases that modify reservations, payments, appointments, records, or any persistent resource, the success criterion cannot end when the agent speaks a response. It must verify the external state: for example, that the correct resource changed exactly once, that a cancellation took effect, or that an uncertain operation was marked for review. When that verification is not possible, the case should be classified as unverified success, not confirmed success.

Minimum fields for each execution

GroupWhat to recordWhy it matters
Voice configurationVariant, transport, turn detection, parameters, and voicePrevents comparisons across different turn policies.
Temporal inputAudio file or identifier, script, event markers, and perturbationsMakes pauses, overlap, and interruptions repeatable.
DelegationBackend, tools, permissions, retries, and timeoutsSeparates the voice layer from later decisions and actions.
External resultInitial state, expected effect, observed effect, and reconciliation evidenceTurns task completion into an auditable check.
DiagnosisFailure labels and correlated tracesHelps attribute the issue to listening, turn-taking, a tool, or the interface.
03

Report four outcome layers separately

The first layer is conversational dynamics. It measures when the system starts speaking, whether it interrupts improperly, whether it lets the user finish, whether it responds to a genuine interruption, and whether it uses overlap in a tolerable way. These phenomena are the focus of Full‑Duplex‑Bench, which proposes an evaluation of turn-taking capabilities in spoken dialogue. They are not equivalent to correctly understanding conversational content.

The second layer is comprehension and fidelity. The question here is whether critical data were captured and retained: names, numbers, dates, negations, alternatives, self-corrections, and constraints. It also matters whether the response reflects the currently active goal. A conversation can sound confident and coherent while changing “Tuesday” into “Thursday” or following an instruction that the user has already corrected.

The third layer is task resolution. Evaluate the decision and tool sequence against a predefined strict criterion. If the task is to find a reservation and modify it, the agent must identify the correct reservation, apply the authorized change, and communicate an outcome consistent with the final state. A satisfactory spoken answer without an effective modification does not meet the success criterion; neither does a correct modification achieved by randomly selecting one ambiguous identity.

The fourth layer is operational safety. It concerns what happens when conditions change: the user interrupts, cancels, corrects an amount, a tool returns late, the connection drops, or an action is no longer relevant. The evaluation should determine whether the operation was stopped, compensated, reconciled, or explicitly declared uncertain. This layer differs from content safety: it concerns control of effects and task state.

04

Build difficult, observable cases

A useful corpus combines nominal scenarios with controlled perturbations. Perturbations are not acoustic decoration: each one should test a hypothesis. A long pause may mean that the user has finished, or that the user is searching for a number. A phrase addressed to someone else is not necessarily an instruction to the agent. A self-correction may invalidate arguments that were already being prepared for a tool.

Include similar numbers, homophones or uncommon names, addresses, alphabetic references, dates, and negative expressions. Create minimal pairs: two audio clips that are identical except for a number, a negation, or the timing of an interruption. If the outcome changes, diagnosis will be more precise than with open-ended conversations that have no clearly expected condition.

Add background noise, volume changes, accents representative of the intended domain, and graded overlap, provided that data handling and speaker representation are authorized. τ‑Voice proposes an evaluation of full-duplex voice agents in real-world domains and considers conditions such as noise, accents, and interruptions alongside verifiable task completion. That combination suggests that an internal test set should not be limited to clean audio.

Do not use only simulated users or only human conversations. The former provide repeatability and a known task state; the latter reveal pragmatic expectations, unexpected formulations, and conversational signals that a script may omit. When human evaluators are used, conceal the evaluated variant, randomize the order, and provide an annotation guide. Their judgment should complement, not replace, the objective verification of task state.

Process for turning an incident into a regression case

  1. 01Describe the user’s goal, the initial state, and the permitted external effect.
  2. 02Reconstruct an authorized audio input or a temporal script that includes the relevant event.
  3. 03Set temporal markers: speech start, pause, possible turn end, interruption, tool submission, and external response.
  4. 04Define expected outcomes for dynamics, critical data, tool calls, and final state.
  5. 05Run the case several times under the same configuration and record variation, traces, and the external outcome.
  6. 06Assign a probable cause only after reviewing the complete sequence; retain the label as a hypothesis if evidence is insufficient.
05

Measure time without reducing it to a single latency

Latency until the agent’s first speech is useful, but it can be misleading. An early response is negative if it cuts off a meaningful pause; a later response may be correct if it avoids acting while the user is self-correcting. For that reason, measure latency to a relevant response relative to an annotated event, not only relative to the last received audio packet.

Record improper cutoffs: occasions where the system begins a response before the user has finished according to the case annotation. Also record ignored interruptions: occasions where the user introduces a stop or change instruction and the system continues speaking or executing the previous goal beyond the accepted policy. Distinguish these events from short overlaps that do not impede communication or cause an incorrect action.

Recovery after barge-in requires a precise definition. Annotate the moment from which an interruption should take effect, the moment spoken output stops, the moment a pending action is invalidated, and the moment an updated response is produced. If the architecture does not make one of these points observable, report it as an observability limitation.

Report distributions and boundary cases, not only averages. High-percentile delay may matter more than the mean in service workflows. Also break results down by scenario type: clean audio, noise, critical number, goal change, low-risk action, and persistent action. Without that breakdown, an improvement in easy cases can conceal a regression in sensitive ones.

Temporal metrics and interpretation rule

MetricReference eventOutcome that should accompany it
Relevant latencyAnnotated end of an instruction or interruptionWhether the response uses the correct goal.
Improper cutoffStart of a pause that was still meaningfulWhether the user had to repeat or repair information.
Ignored interruptionStart of a stop or correction instructionWhether speech, a call, or an obsolete effect continued.
Interruption recoveryMoment when the change should prevailTime to stop, invalidate, and provide a revised response.
Premature silencePause marked as continuationWhether the system requested clarification or started an action too early.
06

Check comprehension, repairs, and tools

For each scenario, identify a small set of critical data and annotate their expected values. Calculate the share of data captured correctly, but do not treat it as a sufficient measure: an error in a date may have a different impact from an error in a minor preference. Keep separate categories for identity, amount, date, destination, consent, cancellation, and safety constraint.

Evaluate confirmations for both content and timing. Repeating a number correctly before an action can reduce ambiguity; asking for confirmation after an operation has been sent does not correct it. The policy should specify which fields require explicit confirmation and which situations require clarification rather than inference. These criteria depend on the workflow and the level of risk accepted by the organization; no universal threshold can be derived from the cited benchmarks.

The conversational repair rate should count whether the system recognizes a discrepancy, requests the appropriate data, incorporates the correction, and completes or safely abandons the task. Do not count an apology followed by the same wrong action as a repair. Also classify whether the repair was initiated by the agent or depended on the user detecting the problem.

At the tool layer, keep a correlation identifier linking the turn, decision, call, and external effect. Check that the sequence is valid and that the arguments correspond to the latest version of the user’s intent. Provider information distinguishes evaluations of conversational dynamics from task-success outcomes and mentions tests of tool-call sequences; this separation is an additional reason not to publish one global score.

07

Combine automation, human review, and instrumented traces

Automate whatever has an observable condition: matching of critical data, call ordering, the presence of a cancellation, reservation state, and differences between expected and observed state. Automation improves coverage and repeatability, but it inherits the limitations of its oracle. If the backend does not expose the final effect or a policy has not been formalized, an automated test can create an unjustified appearance of certainty.

Blinded human review is appropriate for assessing whether a pause was reasonably interpretable, whether an overlap impaired comprehension, or whether a response was pragmatically appropriate. Use at least two reviewers when the impact of the case warrants it, provide operational definitions, and record disagreements. Agreement among reviewers does not turn their judgment into task truth; it helps identify protocol ambiguities and experience-related aspects that traces do not capture.

Instrumented traces are required to attribute failures. They should make it possible to reconstruct, with comparable timestamps, input audio or its protected reference, turn-detection events, transcripts when available, voice output, delegation decisions, tool calls, results, and cancellation state. Establish access controls, retention periods, and minimization procedures consistent with the obligations that apply to recordings and customer data.

Separate development evaluation from release evaluation. During development, a regression set can grow with incidents. Before expanding use, reserve cases that have not guided design decisions. Also repeat executions: systems that combine models, networks, and external services may vary across runs. Report that variation instead of attributing an isolated result to a stable capability.

Evidence triangulation for each case

  1. 01Use an automated checker to validate critical fields, tool sequence, and final state when a reliable oracle exists.
  2. 02Review the interaction blindly to rate turn phenomena and the need for repair.
  3. 03Consult traces to locate turn detection, delegation, cancellation, and external effect in time.
  4. 04If the three sources disagree, do not force a conclusion: label the case for investigation and state what evidence is missing.
  5. 05Aggregate results by layer and scenario, preserving internal links to failure examples and their audit artifacts.
08

Compare benchmarks and provider results carefully

Full‑Duplex‑Bench focuses on turn-taking capabilities in spoken dialogue. τ‑Voice addresses full-duplex voice agents in real-world domains and combines interaction with verifiable task completion. Provider information about GPT‑Live‑1 distinguishes Full Duplex Bench, oriented toward pauses, turns, interruptions, and backchannels, from Tau3 or TauBanking, associated with task success through Pass@1. These labels indicate that the percentages answer different questions.

Before comparing an internal figure with a published one, check the task definition, case set, language, audio type, role of simulated users or human evaluators, number of runs, scoring rule, and success condition. Also check the exact model, evaluation date, voice configuration, turn detection, tools, delegated backend, and permissions. A shared benchmark name does not guarantee experimental equivalence.

Pass@1 should be read precisely: it expresses a first-opportunity success outcome according to the definition of the relevant benchmark, not a general reliability guarantee for every workflow. Nor does it by itself indicate who initiated a repair, how many interruptions were ignored, or whether an obsolete action reached an external system. Use it alongside error breakdowns and evidence of final state.

There are important uncertainties in any move to production. The cited documents do not establish acceptable risk for each organization, do not replace testing on its telephony, integrations, or users, and may not reflect every network or speech variation in its environment. The deployment decision should be based on a representative corpus and an explicit policy for permitted effects.

Comparability checklist before contrasting percentages

DimensionMust match or be declaredRisk if omitted
Evaluated objectiveTurn-taking, comprehension, task, or operational safetyMistaking a fluency metric for a success metric.
ConfigurationModel, delegated backend, tools, and turn detectionAttributing an architectural effect to the model.
Population and audioLanguage, noise, accents, script, simulation, or human interactionGeneralizing from unrepresentative conditions.
ScoringCase unit, number of runs, and success ruleComparing different denominators or criteria.
External effectSystem, permissions, cancellation, and reconciliationDeclaring success without verifying consequences.
09

Turn results into a deployment decision

The final report should present a configuration sheet and four result panels: dynamics, comprehension, task, and operational safety. In each panel, include the set size, scenarios, outcome distribution, serious errors, variation across executions, and observability limits. Add representative examples, both successful and failed, without unnecessarily exposing sensitive content.

Define thresholds by workflow, not by a generic number. In an informational inquiry, moderate delay may be acceptable if the system requests clarification when uncertain. In a reservation change, the priority may be ensuring that corrections prevail and that no action is committed before conditions are met. In an operation with financial or regulatory consequences, an ambiguous critical datum may require escalation or human confirmation. These are policy and risk criteria; they should be approved before observing the result to avoid adjusting the threshold after the fact.

Classify findings as blocking, requiring review, or compatible with a limited canary. Candidates for blocking include incorrect external effects, unconfirmed cancellations, identity or amount errors above the workflow limit, and insufficient traceability to investigate. Candidates for review include pacing problems that do not alter data or actions, provided they do not conceal an accessibility degradation. A canary requires scope limits, monitoring, rollback, and a clear path to human support.

The methodological conclusion is straightforward: GPT‑Live‑1 should be evaluated as part of a voice system, not only as a voice that responds. A test suite that separates listening, turn-taking, response, delegation, and external effect makes it possible to locate problems and decide, based on evidence, what can be expanded. A convincing conversation is an experience signal; on its own, it is not proof that the agent resolved the user’s need correctly and in a controlled manner.

Release decision template

  1. 01Set the workflow, the potential harm, and the external effects that are permitted.
  2. 02Declare separate thresholds for turn-taking, critical data, verifiable success, and cancellation or reconciliation.
  3. 03Run the reserved set and analyze results by perturbation type and configuration.
  4. 04Block expansion when there are erroneous effects, untreated uncertain states, or insufficient traces.
  5. 05If a canary is appropriate, limit the population and actions, monitor the same indicators, and prepare rollback and human escalation.
  6. 06Review thresholds and the corpus when the model, turn detection, backend, tools, or telephony integration changes.

Open questions

  • The provided sources describe benchmarks and general capabilities, but they do not establish universal acceptable-error thresholds for amounts, identities, dates, or cancellations.
  • The available documentation is insufficient to infer that a specific GPT‑Live‑1 configuration will reproduce published results in a different telephony, tool, and user environment.
  • Failure attribution may remain uncertain when temporal traces do not connect audio, decision, tool, and external effect.
  • Retaining and reviewing audio, transcripts, and traces requires legal, contractual, and privacy requirements that are not detailed in the provided sources.
10

Keep exploring

10

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction