M
MISTRAL AI / voxtral-mini-tts-2603

Voxtral TTS

Multilingual text-to-speech model with streaming and zero-shot cloning from a voice reference.

01 / SUMMARY

The essential

ORIENTATIONText-to-speech and cloning

Primary use stated or inferred cautiously from official documentation.

CONTEXTText and voice reference

Maximum input capacity when published by the source.

EXITStreaming audio

Documented maximum generation limit.

ACCESSMistral API and weights

Verified availability channels.

MODEL TYPEVoice and audio

Primary functional family and verified specialties.

LICENSECC BY-NC 4.0

Declared terms for API access, model weights, or self-hosting.

02 / TECHNICAL SHEET

Limits and integration

API IDvoxtral-mini-tts-2603
Model typeVoice and audio
Access modelAPI and open weights
LicenseCC BY-NC 4.0
DeploymentAPI or local with license restrictions
Release23 MAR 2026
Knowledge cutoffNot published
EntranceText · Reference audio
ExitAudio
Context windowText and voice reference
maximum outputStreaming audio
ReasoningNot applicable or not published
Published tools
Structured OutingsNot applicable or not published
Batch processingNot published
Prompt cacheNot published
Fine-tuningDownloadable weights
Verified platformsMistral API
BEST SUITED FOR

Narration, voice prototyping, and multilingual synthesis.

WORTH MONITORING

The license is non-commercial; review consent and voice rights before cloning.

03 / CAPABILITIES

What it can do

01

Orientation

Narration, voice prototyping, and multilingual synthesis.

02

Context

Text and voice reference · Streaming audio

03

Tools and integration

Not published

04

Access

Mistral API and weights

Input modalities
TextReference audio
04 / PRICES

Documented cost

ConceptWorthUnit/condition
Standard inputNot publishedStandard API rate
Cached inputNot publishedReading reused prefixes
Cache write or storageNot publishedThe condition varies by provider
Standard outputNot publishedMay include reasoning tokens
Batch inputNot publishedAsynchronous processing
Batch outputNot publishedAsynchronous processing

Consult the primary source: the billing unit depends on the model type and access channel.

i Prices change and may depend on level, region or context length. Check the source before making a decision.

05 / EVALUATIONS

How to read the results

“

A public benchmark provides guidance, but does not replace an evaluation with your data, tools, budget, and error tolerance.

BenchmarkResultMetricSource
Evaluaciones Voxtral TTS~90 msPublished time to first audioView source ↗

Inferama only highlights a “best result” when the metric, test set, configuration, and date allow for an equivalent comparison. The supplier's figures are presented as claims from its own source.

06 / VERSIONS

Chronology

Voxtral TTS

Version added to Inferama's verified catalog.

ANALYSIS

Related analysis

07 / SOURCES

Traceability