VERIFIED ORGANISATION

DeepSeek

Explore the documented models, published evaluations and primary sources associated with this organisation.

01 / MODELS

Verified catalogue

02 / EVALUATIONS

Published results

Figures are shown with the context reported by their source. A provider result is not an independent comparison and does not replace your own evaluation.

ModelBenchmarkResultMetric
DeepSeek V4.1 FlashGPQA Diamond90,9 %Accuracy
DeepSeek V4.1 FlashTerminal-Bench 2.190,6 %Accuracy
DeepSeek V4.1 FlashDeepSWE v1.174,2 %Resolved
DeepSeek V3.2Evaluaciones DeepSeek V3.2Published GPT‑5 levelReasoning and agents
DeepSeek R1AIME 202479,8 %Published Pass@1
03 / METHOD

How to read the directory

An organization, a product, and a model are not the same unit. Inferama separates them to avoid attributing capabilities or commercial terms to the wrong item.

01

Laboratory

Entity that develops or publishes the model and maintains its technical and safety documentation.

02

Model

Identifiable version with limits, modalities, and behavior that may change between releases.

03

Access channel

API, application, associated cloud, or commercial plan; each channel may have different pricing, retention, and limits.

04

Source and date

Every claim must retain the official page consulted and the verification date.

ANALYSIS

Related analysis

DeepSeek V4.1 Flash: How to Test Whether Its Compressed Cache Reduces an Agent’s Real Cost
ANALISIS

DeepSeek V4.1 Flash: How to Test Whether Its Compressed Cache Reduces an Agent’s Real Cost

DeepSeek says V4.1 Flash has a much smaller persistent KV cache than its predecessor. That figure describes an infrastructure property, not the cost or completion time of an entire task. This analysis proposes a controlled test to measure both.

28 Sep 2026
DeepSeek V3.2 with Tools: How to Audit Reasoning State Before Keeping or Migrating an Agent
ANALISIS

DeepSeek V3.2 with Tools: How to Audit Reasoning State Before Keeping or Migrating an Agent

An operational guide to checking whether an integration preserves the state required across DeepSeek, tools, and the user. It includes reproducible tests, logging criteria, and limits on what can be concluded from `reasoning_content`.

28 Sep 2026
DeepSeek R1 in Production: How to Test the Prompt Template Without Breaking Your Application
ANALISIS

DeepSeek R1 in Production: How to Test the Prompt Template Without Breaking Your Application

Official prompting recommendations for DeepSeek R1 are hypotheses worth validating in each application. This protocol compares instruction placement and the “<think>” prefix while measuring both answer quality and integration failures.

27 Sep 2026
DeepSeek V4.1 Flash + Eleven v3 vs. full-duplex voice: how to compare the architectures
COMPARATIVA

DeepSeek V4.1 Flash + Eleven v3 vs. full-duplex voice: how to compare the architectures

A modular chain built from speech recognition, DeepSeek V4.1 Flash, and Eleven v3 is not equivalent to a full-duplex voice model. This guide proposes comparing both systems end to end, without declaring a winner based on isolated specifications.

26 Sep 2026
DeepSWE v1.1: What a Patch That Passes Its Verifier Measures—and When Its pass@1 Stops Showing That the Agent Solved the Task
ANALISIS

DeepSWE v1.1: What a Patch That Passes Its Verifier Measures—and When Its pass@1 Stops Showing That the Agent Solved the Task

DeepSWE v1.1 evaluates code changes produced by agents on original, long-horizon tasks. Its result can provide a useful signal, but a leaderboard row is interpretable only when the harness, error policy, rollouts, verifier, and exact benchmark version are known. A review by Epoch AI adds a specific warning: under certain evaluation conditions, some patches that modify tests can receive a false negative.

23 Sep 2026
DeepSeek: how to choose between the official API, published weights, and external providers without confusing openness with operational control
ANALISIS

DeepSeek: how to choose between the official API, published weights, and external providers without confusing openness with operational control

The availability of weights, an API-compatible endpoint, or a family name does not by itself settle a production decision. This analysis separates the artifact, license, endpoint, operator, and evidence needed to assess DeepSeek through the official API, self-managed infrastructure, or an external provider.

22 Sep 2026
04 / SOURCES

Traceability