Primary use stated or inferred cautiously from official documentation.
DeepSeek V4.1 Flash
Agents, coding, and open deployment at low cost.
The essential
Maximum input capacity when published by the source.
Documented maximum generation limit.
Verified availability channels.
Primary functional family and verified specialties.
Declared terms for API access, model weights, or self-hosting.
Limits and integration
| API ID | deepseek-flash |
|---|---|
| Model type | Text and reasoning |
| Access model | Paid proprietary |
| License | Proprietary |
| Deployment | Hosted API |
| Release | 10 SEP 2026 |
| Knowledge cutoff | Not published |
| Entrance | Text · Image |
| Exit | Text |
| Context window | 1,000,000 tokens |
| maximum output | 384.000 tokens |
| Reasoning | Modes with and without reasoning |
| Published tools | Features · Responses API · Chat prefix · FIM |
| Structured Outings | Supported |
| Batch processing | Not published |
| Prompt cache | Supported |
| Fine-tuning | Open weights |
| Verified platforms | DeepSeek API · Hugging Face |
Agents, coding, and open deployment at low cost.
Execution conditions vary across benchmarks; do not compare the number alone.
What it can do
Orientation
Agents, coding, and open deployment at low cost.
Context
1.000.000 tokens · 384.000 tokens
Tools and integration
Features · Responses API · Chat prefix · FIM
Access
DeepSeek API · open weights
Documented cost
| Concept | Worth | Unit/condition |
|---|---|---|
| Standard input | 0,30 USD / 1 M tokens (peak) | Standard API rate |
| Cached input | 0,006 USD / 1 M tokens (peak) | Reading reused prefixes |
| Cache write or storage | Not published | The condition varies by provider |
| Standard output | 1,20 USD / 1 M tokens (peak) | May include reasoning tokens |
| Batch input | Not published | Asynchronous processing |
| Batch output | Not published | Asynchronous processing |
Off-peak rates are half price; schedules are published in UTC.
i Prices change and may depend on level, region or context length. Check the source before making a decision.
How to read the results
A public benchmark provides guidance, but does not replace an evaluation with your data, tools, budget, and error tolerance.
| Benchmark | Result | Metric | Source |
|---|---|---|---|
| GPQA Diamond | 90,9 % | Accuracy | View source ↗ |
| Terminal-Bench 2.1 | 90,6 % | Accuracy | View source ↗ |
| DeepSWE v1.1 | 74,2 % | Resolved | View source ↗ |
Inferama only highlights a “best result” when the metric, test set, configuration, and date allow for an equivalent comparison. The supplier's figures are presented as claims from its own source.
Chronology
DeepSeek V4.1 Flash
Version added to Inferama's verified catalog.
Related analysis
DeepSeek V4.1 Flash: How to Test Whether Its Compressed Cache Reduces an Agent’s Real Cost
DeepSeek says V4.1 Flash has a much smaller persistent KV cache than its predecessor. That figure describes an infrastructure property, not the cost or completion time of an entire task. This analysis proposes a controlled test to measure both.
28 Sep 2026 ↗ COMPARATIVADeepSeek V4.1 Flash + Eleven v3 vs. full-duplex voice: how to compare the architectures
A modular chain built from speech recognition, DeepSeek V4.1 Flash, and Eleven v3 is not equivalent to a full-duplex voice model. This guide proposes comparing both systems end to end, without declaring a winner based on isolated specifications.
26 Sep 2026 ↗ ANALISISDeepSWE v1.1: What a Patch That Passes Its Verifier Measures—and When Its pass@1 Stops Showing That the Agent Solved the Task
DeepSWE v1.1 evaluates code changes produced by agents on original, long-horizon tasks. Its result can provide a useful signal, but a leaderboard row is interpretable only when the harness, error policy, rollouts, verifier, and exact benchmark version are known. A review by Epoch AI adds a specific warning: under certain evaluation conditions, some patches that modify tests can receive a false negative.
23 Sep 2026 ↗