Primary use stated or inferred cautiously from official documentation.
Scribe v2
Speech recognition model for more than 90 languages, with diarization, word-level timestamps, entities, and audio labels.
The essential
Maximum input capacity when published by the source.
Documented maximum generation limit.
Verified availability channels.
Primary functional family and verified specialties.
Declared terms for API access, model weights, or self-hosting.
Limits and integration
| API ID | scribe_v2 |
|---|---|
| Model type | Voice and audio |
| Access model | Paid proprietary |
| License | Proprietary |
| Deployment | Hosted API |
| Release | 2026 |
| Knowledge cutoff | Not published |
| Entrance | Audio · Video |
| Exit | Text · Timestamps · Diarization |
| Context window | Audio or video file |
| maximum output | Structured transcription |
| Reasoning | Not applicable or not published |
| Published tools | |
| Structured Outings | Not applicable or not published |
| Batch processing | Not published |
| Prompt cache | Not published |
| Fine-tuning | Not published |
| Verified platforms | ElevenLabs API |
Transcription of meetings, media, and multilingual content.
Measure error by language, number of speakers, noise, and domain terms.
What it can do
Orientation
Transcription of meetings, media, and multilingual content.
Context
Audio or video file · Structured transcription
Tools and integration
Not published
Access
ElevenLabs API
Documented cost
| Concept | Worth | Unit/condition |
|---|---|---|
| Standard input | Not published | Standard API rate |
| Cached input | Not published | Reading reused prefixes |
| Cache write or storage | Not published | The condition varies by provider |
| Standard output | Not published | May include reasoning tokens |
| Batch input | Not published | Asynchronous processing |
| Batch output | Not published | Asynchronous processing |
Consult the primary source: the billing unit depends on the model type and access channel.
i Prices change and may depend on level, region or context length. Check the source before making a decision.
How to read the results
A public benchmark provides guidance, but does not replace an evaluation with your data, tools, budget, and error tolerance.
| Benchmark | Result | Metric | Source |
|---|---|---|---|
| Evaluación Scribe v2 | 90+ languages | Published ASR coverage | View source ↗ |
Inferama only highlights a “best result” when the metric, test set, configuration, and date allow for an equivalent comparison. The supplier's figures are presented as claims from its own source.
Chronology
Scribe v2
Version added to Inferama's verified catalog.
Related analysis
Scribe v2 in speech benchmarks: what its scores compare—and what is still needed to reproduce them
A WER score alone does not describe a transcription system’s overall performance. We examine what can be verified about Scribe v2, how the corpus, normalization, and configuration affect results, and what information is needed to repeat a comparison.
28 Sep 2026 ↗ NOTICIAGPT-Transcribe: what a team should revalidate before replacing its transcription system
GPT-Transcribe is listed as a transcription model available through the OpenAI API. Before replacing an existing ASR system, teams should validate not only the resulting text, but also the technical contract that supports search, summaries, alerts, quotations, reviews, and compliance records.
22 Sep 2026 ↗ ANALISISElevenLabs: How to Separate the Model, Voice, Access Channel, and Data Retention Before Taking Generative Audio to Production
Adopting ElevenLabs for generative audio is not a single decision: every workflow combines a model, a voice asset, a processing channel, and a different retention regime. This guide provides an operational matrix for documenting them, limiting assumptions, and preparing verifiable migrations, access controls, and deletions.
22 Sep 2026 ↗