What Is Being Evaluated—and What Is Not
Evaluating a text-to-speech system is not a matter of deciding whether one sample “sounds better” than another in absolute terms. For a product, localization, or audiovisual production team, the useful question is more specific: given a voice, a script, and a delivery cue, what changes in the audio, and how consistently? A test focused on audio tags in Eleven v3 should measure text fidelity, perceived expressive delivery, voice continuity, and consistency across generations as separate dimensions.
This protocol does not offer a general naturalness ranking or compare providers. Nor does it assume that a tag will always produce a particular emotion. The documentation provided for this article describes generation parameters and voice settings, but does not make it possible to verify which specific tags are compatible with Eleven v3 or what effects they are expected to have. Before generating audio, therefore, the team should confirm in the current documentation for the access method it plans to use which cues the model supports and how they should be written. Without that check, tags are an experimental hypothesis, not a feature confirmed by this article.
The distinction matters: the fact that an instruction appears in a script does not prove that the system will interpret it consistently. The aim is to establish what evidence supports a decision to use a particular voice for a particular task, and which questions remain unanswered. Findings apply to the conditions actually tested, not to every voice, language, or version.
Preparation: Fix the Conditions Before Listening
Start by defining the intended use. Informational narration, customer-service messages, and dramatized dialogue require different behavior: emphatic intonation may suit a character but be inappropriate for a service instruction. Do not combine these tasks into a single score. Prepare a small set of representative scripts for each category, including short and long sentences, varied punctuation, and vocabulary relevant to the product. Avoid text that contains personal data or that you are not authorized to process.
For every generation, record the date, language, access method, available model and voice identifiers, exact text, and all settings. ElevenLabs’ speech-generation reference includes fields for model and voice identification, voice settings, and a seed; record the parameters you actually used, without assuming that every access method exposes the same options. The model list is a reference point for identifying available models and capabilities, but the record of the specific generation remains essential.
Choose authorized voices and keep the voice constant within each comparison. Also keep the text, language, and all other parameters identical when changing one experimental condition. If you change the tag, stability setting, and script at the same time, you will not be able to attribute the difference clearly to any one cause. The design can include multiple voices and languages, but analyze their results separately.
Decide in advance what it means for the requested delivery to have been achieved. For an instruction such as “calm tone,” specify which audible characteristics would count as calm in that context, without requiring every listener to describe the emotion using the same word. Defining this beforehand reduces the risk of changing the criteria to fit a result already heard.
Preparation sequence
- 01Define the task and select representative scripts for each category.
- 02Check current documentation for the cues accepted by the model and access method.
- 03Use authorized voices and record the language, date, available version or identifier, and parameters.
- 04Define in advance which audio characteristics would count as meeting each delivery goal.
- 05Generate the conditions using the same text, changing only the variable being studied.
Designing the Conditions: Baseline, Cues, and Stability
An initial comparison can include three types of condition: text with no added expressive cue; text with a cue whose compatibility has been confirmed; and text with that same cue under other stability settings that the access method allows you to control. Document the exact tag or instruction, where it appears in the text, and why you expect to hear the specified delivery. Do not invent a syntax or treat a phrase written in the script as a control tag unless the access method’s documentation confirms it.
Stability settings are not equivalent to a measure of quality. ElevenLabs documentation describes stability and voice similarity as settings, and relates stability to randomness and emotional range. That is a reason to test the parameter as a variable, but it does not let you conclude in advance which value will be better for a task. Rather than looking for the highest or lowest number, compare documented configurations and observe the trade-off between expressiveness, consistency, and fidelity.
Include multiple generations for each condition. A single sample may reflect a one-off variation rather than typical behavior. If the access method offers a seed, record its value and decide how you will use it: keeping it fixed may help structure comparisons, while generating with different values may help you study variation. Do not assume that a seed guarantees an identical repeat; that property must be checked for the specific system and configuration.
Keep the script’s length and content stable across conditions. If a tag requires changing the text, create a separate, explicit comparison, because you will no longer be isolating the effect of the cue. Keep the original file for every result along with a readable record of its parameters, subject to the applicable storage and usage rules.
Basic comparison matrix
| Condition | What changes | What it helps you observe |
|---|---|---|
| Baseline | No added cue; parameters recorded | How the script is interpreted without the intervention being studied |
| Expressive cue | A previously verified tag or instruction is added | Whether perceived delivery is close to the defined goal |
| Stability variation | A documented setting changes; everything else stays constant | How expressiveness, voice continuity, and repeatability vary |
| Repetitions | The same condition is generated again | How often a particular result or failure occurs |
Three Tasks, Three Listening Criteria
For informational narration, prioritize comprehension and a level of neutrality that suits the material. Listen for whether the prosody highlights the relationships between ideas without dramatizing sentences that call for a restrained reading. A passage with dates, acronyms, or figures may reveal intelligibility problems that do not appear in a simple sentence. The question is not whether the narration is emotional, but whether its delivery suits the purpose and preserves the content.
For customer service, check whether the voice communicates courtesy and clarity without sounding mechanical or overly familiar. Include messages with instructions, confirmations, and potentially frustrating situations. An expressive cue may sound convincing in a greeting yet fail in an apology or safety instruction. Evaluate each subcase as part of the same task, but keep its results separate.
In dramatized dialogue, more pronounced variation may be part of the goal. Evaluate whether the lines are differentiated as the script calls for and whether the voice remains recognizable as the emotion or intensity changes. Voice identity continuity does not mean that every line should sound the same: it means interpretive changes should not make it seem as though the speaker was accidentally replaced. Clarify what variation is desirable before scoring.
For all tasks, avoid letting evaluators know the condition in advance. Present anonymized files in random order at comparable playback levels. Recruit listeners who are competent in the language of the sample and record their familiarity with the type of content. Anonymization does not remove every bias, but it reduces the influence of expectations about a tag or setting.
Listening Rubric and Supporting Measures
Score four dimensions separately. The first is lexical fidelity: whether words were omitted, added, or changed. The second is delivery compliance: whether listeners perceived the characteristics defined as the goal, not whether they necessarily guessed which tag was used. The third is continuity of vocal identity across the set of samples. The fourth covers audible problems such as artifacts, dropouts, inappropriate pauses, or volume changes that affect usability.
Use short scales with explicit descriptors and an option for “cannot determine.” Also ask for a brief note when a score is low or evaluators disagree. Do not reduce the result to a single average: an average can hide the fact that a voice works well for narration and poorly for dialogue, or that two listeners disagree systematically. Report the distribution of scores and the frequency of each type of failure.
Automatic transcription can help locate possible differences between the intended text and what was recognized, but it is not a conclusive measure of intelligibility. Recognition errors may be transcription errors rather than audio errors; conversely, a correct transcription does not guarantee that a voice sounds clear or appropriate. Check relevant discrepancies against human annotation before classifying them as lexical errors by the system.
Duration, pauses, and variation between generations are supporting measures. They can describe differences and flag cases for review, but they do not replace listening. Speech-synthesis evaluation also distinguishes dimensions such as quality, speaker similarity, and intelligibility; treating them separately helps prevent a single impression from absorbing different criteria. Do not extrapolate a measure obtained for one task or language to other tasks or languages.
Minimum per-sample rubric
| Dimension | Evaluation question | Recommended record |
|---|---|---|
| Lexical fidelity | Does the audio match the intended text? | Omissions, substitutions, additions, and cases to review |
| Perceived delivery | Are the characteristics defined before the test audible? | Score by listener and a brief description |
| Voice identity | Does the voice remain recognizable across the set? | Separate assessment of expected expressive changes |
| Artifacts and pauses | Are there failures that hinder comprehension or use? | Type, location, and frequency of the problem |
Repeatability, Analysis, and Reporting Results
Define in advance how many generations you will make for each condition and what counts as a failure. The number should be sufficient for the team to observe variation within its available resources, and it should be reported with the results. The sources provided do not establish a universal threshold for declaring a configuration reliable for every voice or task. If a decision has significant consequences, treat a small sample as exploratory and use it to motivate a larger test, not a categorical conclusion.
Summarize the data by voice, language, task, and condition. Present both the center of the distribution and its spread; show disagreement between listeners and examples of relevant failures. If a cue improves perceived delivery but also increases omissions, artifacts, or variation, make that trade-off visible. Also report cases where the cue did not produce the expected effect: excluding them would give an incomplete picture of predictability.
Automatic transcripts, duration, and other computable measurements can accompany the analysis, provided you describe how they were obtained and what their limitations are. Do not combine their results with human ratings as if they measured the same thing. If the model, access method, voice, or parameters change, record the change and consider whether the evaluation needs to be repeated to maintain comparability.
The final report should allow another person to reproduce the design: scripts, language, authorized voices, identifiers, parameters, conditions, repetitions, instructions for listeners, and scoring rules. You do not need to publish audio for the analysis to be useful. If you decide to share samples, separately verify the applicable terms and voice permissions, along with any conditions that affect publication.
Interpreting an adoption decision
- 01First check that the conditions and files are complete and correctly identified.
- 02Review lexical fidelity and failures that could prevent the intended use.
- 03Compare perceived delivery with the criteria defined before listening.
- 04Examine voice continuity, variation between generations, and disagreements between listeners.
- 05Adopt the configuration only for the task, voice, and language covered by the evidence; expand the test if relevant cases are missing.
Limitations and Practical Criteria
A controlled test does not eliminate sensitivity to the script, language, voice, or service changes. Findings apply to the available version or identifier, access method, and settings recorded at the time of generation. If any of these elements change, the previous evidence may not describe the new behavior. Record the execution date and avoid turning a local result into a claim about Eleven v3 in general.
There are also limits to human interpretation. Listeners may disagree about whether a voice sounds calm, convincing, or approachable. That disagreement is a result to report, not noise to erase with an average. If people do not identify a delivery consistently, the team can review the instruction, the definition of success, or the use case, and design an additional test.
For a decision, require evidence aligned with the risk: acceptable lexical fidelity for the content, delivery appropriate to the task, sufficient voice continuity, and variation compatible with the production workflow. If the configuration fails on a critical dimension, a high score on another should not conceal it. If there are few samples, listeners disagree, or tag compatibility was not verified, the appropriate conclusion is provisional.
This turns an evaluation of Eleven v3 into a local, auditable protocol rather than a promise of universal results. The team can document which combinations appear to work, where variation occurs, and what needs to be tested before expanding use. Evaluation does not replace permission checks or editorial oversight; it provides evidence for decisions with explicit limitations.
Open questions
- The verified sources available do not document which specific audio tags Eleven v3 supports or their syntax; this must be checked in the current documentation for the access method.
- No universal number of repetitions, listener threshold, or minimum adoption score has been established; the protocol should be adapted to the risk and purpose of each evaluation.
- The behavior of one voice, language, or configuration does not automatically predict the behavior of other voices, languages, versions, or access methods.
- Available information about the parameters does not guarantee that a seed will produce an identical output or that settings are exposed in the same way across all access methods.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction