Ilustración editorial para Scribe v2 en los benchmarks de voz: qué comparan sus puntuaciones y qué falta para reproducirlas
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

A leadership claim needs a protocol behind it

ElevenLabs introduced Scribe v2 with the claim that it achieved the lowest WER in industry benchmarks. That is a claim made by the provider, not an independent conclusion that can be accepted without knowing which tests were compared, which versions were used, and what scoring rules applied. The launch page lets readers attribute the claim correctly, but does not identify the specific benchmarks supporting it or provide enough of a protocol to repeat the calculation.

This does not show that the claim is false. It means that, based on the information documented on that page, an outside reader cannot reconstruct which audio was used, how large the sample was, how the reference transcripts were prepared, or which other systems took part. Without those details, “the lowest WER” should be read as a statement whose scope is not established publicly by that source—not as proof of universal superiority.

There is also an independent evaluation that should be kept separate from the launch claim: Artificial Analysis publishes AA-WER v2.0, a benchmark that reports results for Scribe v2. Its existence offers a concrete reference to examine, but it does not automatically verify the benchmarks ElevenLabs alludes to. Each result belongs to its own dataset, protocol, and execution date.

02

What WER measures—and what it leaves out

WER, or word error rate, compares an automatic transcript with a reference transcript. It counts substitutions, deletions, and insertions relative to the reference text, and expresses their total as a proportion of the words in that text. It is useful for summarizing discrepancies, but it does not, by itself, measure how useful a transcript is or how serious each error may be in practice.

Two systems can receive the same score and make different kinds of mistakes. Misrecognizing a proper name may be more costly in a meeting than omitting a filler word; however, an aggregate metric does not assign importance according to context. Nor does WER directly report latency, the stability of live output, diarization, readability, or cost. Those dimensions require additional measures and tests.

The result also depends on what counts as a word and on how both the reference and the system output are normalized before comparison. Differences in capitalization, punctuation, numbers, contractions, or disfluencies can change the count if the rules are not shared. NIST documents reference normalization as part of its OpenASR evaluation plan; publishing the rule is therefore part of the protocol, not an editorial detail.

The unit of comparison matters too. A reference may divide speech differently from the model’s transcript. If a system works in chunks, if segments are concatenated, or if portions of the audio are excluded, the procedure can change which errors enter the calculation. A number without a defined corpus, reference, normalization, and segmentation method therefore does not make clear exactly what is being measured.

03

The corpus defines the scope of the conclusion

A score summarizes performance on a particular dataset, not on every kind of audio an organization might process. The selection of recordings, their duration, languages, accents, acoustic conditions, and speech types determines which population of examples is represented. An average can hide differences between languages or conditions if results are not reported by group along with the number of samples in each group.

FLEURS provides context for one kind of multilingual evaluation: the original paper describes a dataset covering 102 languages and explains its purpose in evaluating speech representations. The fact that a benchmark includes several languages does not mean it comprehensively measures every real-world use case, or that an average across languages reflects the needs of every team. Interpreting a result requires knowing which languages were evaluated, how they were weighted, and how many samples contributed to each score.

AA-WER v2.0 identifies several datasets in its evaluation and distinguishes AA-AgentTalk, which is proprietary, from VoxPopuli and Earnings22. This distinction matters: a benchmark may combine datasets with different conditions and levels of availability. Including datasets with restricted access can make it harder for others to reproduce the complete evaluation identically, even if they can examine the published protocol or repeat parts using available data.

Audio licensing also affects reproducibility. A dataset may be well known and described without necessarily being freely redistributable or available for repeating a test. If the files cannot be shared, the documentation should explain how access was obtained, which subset was used, and what alternative artifacts are available to make the comparison auditable.

Questions that help define the scope of a result

ElementWhat to checkWhy it affects interpretation
CorpusName, version, license, and inclusion criteriaDefines which data and use cases the test represents
SampleNumber of recordings, duration, and exclusionsHelps assess coverage and possible selection bias
Languages and conditionsResults by group and the size of each groupPrevents an average from hiding important differences
WeightingHow results from each dataset are combinedAn average can change depending on the weight assigned to each source
04

Configuration, prompting, and product variants

A model label alone is not enough to identify a reproducible run. Comparing a score requires, at a minimum, the exact model identifier and the access or execution date, along with relevant parameters. Services can be updated, and an evaluation run on one date does not necessarily represent the behavior of a later version. Available capability documentation helps distinguish products, but does not, by itself, reveal the historical configuration used in a launch evaluation.

It is also necessary to clarify whether keyterm prompting was used. ElevenLabs documentation describes the keyterms parameter for its batch API. Providing known terms can help guide recognition of specialized vocabulary; therefore, an evaluation with that information is not directly equivalent to one without it. To interpret the result, the terms provided, the selection criteria, and whether all compared systems received the same kind of contextual advantage should be published.

Scribe v2 and Scribe v2 Realtime should not be treated as a single benchmark entry without specifying which product was evaluated. ElevenLabs distinguishes them in its transcription offering. A test that processes recorded audio in batch mode and a test of live transcription address different conditions: in the latter, progressive audio input and latency may matter alongside final accuracy. If results from the two modes are mixed, a single ranking loses its meaning.

Artificial Analysis’s methodology also distinguishes batch evaluation from streaming evaluation. That distinction is a practical reason to require every score to specify the mode used. It does not, by itself, let us infer which mode was used for each claim on the ElevenLabs launch page; that information would need to accompany the specific score.

Configuration information that should accompany a score

DetailAudit questionRisk if missing
Model and versionWhich exact identifier and execution date were recorded?It is impossible to know whether someone else evaluated the same version
ModeWas recorded audio processed, or was live transcription evaluated?Different functional conditions are being compared
KeytermsWere terms provided, and which ones?A contextual advantage may be mistaken for unaided performance
SegmentationHow was audio divided, concatenated, or excluded?The unit being scored may not be equivalent across systems
05

When two scores are comparable

The strongest comparison uses the same audio files, the same references, and the same scoring rules. It also keeps language, segmentation, and access to context equivalent. If providers receive different instructions or vocabularies, or one system processes chunks while another receives complete recordings, the result may reflect those differences as well as recognition ability.

In practice, it is not always possible to run every service under perfectly identical conditions. In that case, the right approach is to document the differences and limit the conclusion. A benchmark can still guide an evaluation even if it does not support a strict causal comparison between models. The problem arises when a score is presented without the conditions needed to distinguish system performance from protocol effects.

Artificial Analysis publishes a methodology for its speech-to-text benchmark and a comparative non-streaming table. These sources make it possible to place Scribe v2’s score within a specific evaluation. They do not automatically turn its position there into a claim about every benchmark, language, audio type, or real-time variant. A position in a table answers a question defined by that table’s rules and participants.

A short process for auditing a comparison

  1. 01Find the original claim and note which product, metric, and scope it refers to.
  2. 02Identify the benchmark, its version, the included datasets, their licenses, and the sample size.
  3. 03Review the references, normalization, segmentation, and scoring formula.
  4. 04Record the model identifier, execution date, batch or real-time mode, and whether keyterms were used.
  5. 05Check whether all systems received the same audio and conditions; separate results when they did not.
  6. 06Limit the conclusion to the datasets, languages, and modes actually evaluated, and state what cannot be reproduced.
06

What can be reproduced—and what remains unclear

The identified public sources make it possible to describe the launch statement as an ElevenLabs claim and to verify that AA-WER v2.0 is an evaluation that reports a result for Scribe v2. They also make it possible to consult FLEURS’s design, the normalization covered by NIST, Artificial Analysis’s methodology, and ElevenLabs documentation about keyterms and transcription models. These are useful pieces of evidence, but they do not form a single protocol that reproduces every performance claim.

The main uncertainty is which specific benchmarks the launch page meant when it attributed the lowest WER to Scribe v2. The supplied source does not identify those benchmarks or set out their protocols. It also does not establish the exact model version that was run, whether prompting was used, what normalization was applied, or how segments were handled. It would not be rigorous to fill those gaps by assuming the configuration matched a later evaluation.

The existence of an independent comparison table does not resolve every question either. Repeating a benchmark in full requires its artifacts and test conditions, and the availability of proprietary datasets may limit third-party replication. If the evaluator reports aggregate results, readers should check whether results are also published by dataset or condition and whether the weighting is explained; an average is not a substitute for that information.

07

A minimum publication checklist for an ASR benchmark

A benchmark useful to others should make it possible to understand both the number and the process that produced it. Publishing only a final ranking is not enough: readers need to determine whether the corpus represents their use case, whether the rules are consistent, and whether a repeat evaluation is feasible. The following checklist cannot guarantee that two services operate under identical conditions in every respect, but it makes visible the differences that limit comparison.

Uncertainty should also be communicated. Results depend on the sample and may vary across subsets. Publishing the number of examples and disaggregated results helps readers assess that variation; if a margin or interval is reported, the calculation should be explained. Without this information, a small difference between two scores should not automatically become a firm conclusion about which system is better.

What an evaluation report should include

  1. 01Benchmark name and version, datasets, licenses, languages, and selection criteria.
  2. 02A list of files or sample identifiers, total duration, exclusions, and reasons for exclusion, subject to applicable licenses.
  3. 03Reference transcripts or a reproducible description of how they were obtained, together with normalization rules.
  4. 04Model identifier and execution date, usage mode, relevant parameters, and any keyterms provided.
  5. 05Rules for segmentation, concatenation, scoring, and WER calculation.
  6. 06Results by dataset, language, and condition, as well as the aggregate, with sample sizes and weights.
  7. 07Artifacts or instructions for repeating the evaluation, plus an explanation of any access limitations for proprietary data.
08

Conclusion: use the score as a signal, not a guarantee

Published scores can help identify candidates for an organization’s own test, but they do not guarantee performance on a particular organization’s data or audio. The ElevenLabs claim about the lowest WER should remain attributed to the provider, and its scope should not be broadened without identifying the benchmarks and protocols to which it refers. AA-WER v2.0 provides an identifiable evaluation of Scribe v2, but its conclusions belong to that benchmark and its conditions.

Before deciding whether to migrate, a team should test representative samples of its languages, recordings, proper names, and acoustic conditions, apply consistent reference rules, and evaluate separately the features that matter to its workflow. If live transcription is required, the team should evaluate the variant and metrics appropriate to that mode rather than extrapolating from a batch test.

The cautious conclusion is neither that Scribe v2 is a universal winner nor that the reported figures have no value. It is that WER provides a rigorous signal only when accompanied by the corpus, version, configuration, and calculation rules. If those details are missing, the score can serve as an indication, but not as a reproducible comparison or a guarantee of future results.

Open questions

  • The ElevenLabs launch page does not specify which benchmarks support its claim of the lowest WER.
  • The supplied sources do not make it possible to reconstruct the exact version, execution date, configuration, normalization, or segmentation used in the launch tests.
  • The availability of proprietary datasets, such as AA-AgentTalk, may limit full third-party reproduction of the benchmark.
  • There is not enough information to determine whether the launch claim used keyterm prompting or which terms, if any, were provided.
09

Keep exploring

09

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction