Primary use stated or inferred cautiously from official documentation.
Qwen3.8-Omni-Flash-Realtime
Audio-visual model for real-time interaction with voice and text output, tools, and remote MCP.
The essential
Maximum input capacity when published by the source.
Documented maximum generation limit.
Verified availability channels.
Primary functional family and verified specialties.
Declared terms for API access, model weights, or self-hosting.
Limits and integration
| API ID | qwen3.8-omni-flash-realtime |
|---|---|
| Model type | Voice and audio · Multimodal |
| Access model | Paid proprietary |
| License | Proprietary |
| Deployment | Hosted API |
| Release | 21 SEP 2026 |
| Knowledge cutoff | Not published |
| Entrance | Text · Image · Audio · Video |
| Exit | Text · Audio |
| Context window | 196.608 input tokens |
| maximum output | Not published |
| Reasoning | Not published |
| Published tools | Features · Remote MCP · WebSocket · WebRTC |
| Structured Outings | Not published |
| Batch processing | Not published |
| Prompt cache | Not published |
| Fine-tuning | Not published |
| Verified platforms | Alibaba Cloud Model Studio |
Multimodal assistants with video and spoken conversation.
A WebSocket session has a maximum duration; manage turns and accumulated context.
What it can do
Orientation
Multimodal assistants with video and spoken conversation.
Context
196.608 input tokens · Not published
Tools and integration
Functions · Remote MCP · WebSocket · WebRTC
Access
Alibaba Cloud Model Studio
Documented cost
| Concept | Worth | Unit/condition |
|---|---|---|
| Standard input | Not published | Standard API rate |
| Cached input | Not published | Reading reused prefixes |
| Cache write or storage | Not published | The condition varies by provider |
| Standard output | Not published | May include reasoning tokens |
| Batch input | Not published | Asynchronous processing |
| Batch output | Not published | Asynchronous processing |
Consult the primary source before budgeting for a deployment.
i Prices change and may depend on level, region or context length. Check the source before making a decision.
How to read the results
A public benchmark provides guidance, but does not replace an evaluation with your data, tools, budget, and error tolerance.
i We have not found published benchmarks for this model.
Inferama only highlights a “best result” when the metric, test set, configuration, and date allow for an equivalent comparison. The supplier's figures are presented as claims from its own source.
Chronology
Qwen3.8-Omni-Flash-Realtime
Version added to Inferama's verified catalog.