VERIFIED ORGANISATION

OpenAI

Explore the documented models, published evaluations and primary sources associated with this organisation.

01 / MODELS

Verified catalogue

GENERALIST FRONTIER

GPT‑6 Astra

Generalist frontier model oriented towards complex tasks, use of tools and highly demanding professional work.

16 SEP 2026
GENERALIST FRONTIER

GPT‑5.6 Sol

Complex work, coding, and tool-using agents.

16 SEP 2026
COST/CAPABILITY BALANCE

GPT‑5.6 Terra

Tasks requiring a balance of quality, speed, and cost.

16 SEP 2026
HIGH EFFICIENCY

GPT‑5.6 Luna

High volume, subagents, and structured extraction.

16 SEP 2026
PREVIOUS FRONTIER

GPT‑5.5

Previous OpenAI generation for code and professional work, retained for comparing migrations and cost.

18 SEP 2026
NON-REASONING · LONG CONTEXT

GPT‑4.1

Long-context non-reasoning model, useful as a historical reference for code, instruction following, and predictable latency.

18 SEP 2026
OPEN REASONING

gpt‑oss‑120b

Reasoning model with open weights for local or private execution; fits on one H100 GPU according to OpenAI.

18 SEP 2026
VISUAL GENERATION AND EDITING

GPT‑Image‑2.5 Sunburst

OpenAI's main model for generating and editing images with complex instructions and high-quality visual production.

18 SEP 2026
VIDEO WITH SYNCHRONIZED AUDIO

Sora 2 Pro

Historical video model with synchronized audio, useful for comparing the evolution of the Sora family.

18 SEP 2026
REAL-TIME VOICE CONVERSATION

GPT‑Live 1

Real-time voice model for expressive conversations, natural interruptions, and tool-enabled workflows.

18 SEP 2026
SPEECH RECOGNITION

GPT‑Transcribe

Specialized high-accuracy transcription model for files and real-time audio inputs.

18 SEP 2026
TEXT VECTOR REPRESENTATION

text‑embedding‑3‑large

OpenAI's highest-capacity embedding model for semantic search, classification, clustering, and RAG.

18 SEP 2026
CODING, COMPUTER USE AND PROFESSIONAL WORK

GPT‑6.1 Sol

OpenAI model for complex coding, computer use and professional work with text and image inputs. The provider describes performance close to Astra at a lower cost; assess that trade-off on your own tasks.

30 SEP 2026
02 / EVALUATIONS

Published results

Figures are shown with the context reported by their source. A provider result is not an independent comparison and does not replace your own evaluation.

ModelBenchmarkResultMetric
GPT‑6 AstraTerminal-Bench Science 0.164,6 %Accuracy
GPT‑6 AstraGPQA Diamond96,0 %Accuracy
GPT‑6 AstraBenchCAD95,9 %Geometric overlap
GPT‑5.6 SolAgents' Last Exam52,7 %Accuracy
GPT‑5.6 SolSWE-Bench Pro64,6 %Resolved
GPT‑5.6 SolTerminal-Bench 2.188,8 %Accuracy
GPT‑5.6 TerraAgents' Last Exam50,4 %Accuracy
GPT‑5.6 TerraSWE-Bench Pro63,4 %Resolved
GPT‑5.6 TerraTerminal-Bench 2.187,4 %Accuracy
GPT‑5.6 LunaAgents' Last Exam50,3 %Accuracy
GPT‑5.6 LunaSWE-Bench Pro62,7 %Resolved
GPT‑5.6 LunaTerminal-Bench 2.184,7 %Accuracy
GPT‑5.5Evaluaciones oficiales GPT‑5.5PublishedCode and professional work
GPT‑4.1SWE-bench Verified54,6 %Published solved problems
gpt‑oss‑120bEvaluaciones gpt‑ossPublished o4-mini levelGeneral reasoning
GPT‑Image‑2.5 SunburstEvaluación visual GPT‑Image‑2.5Main modelPreference and instruction following
Sora 2 ProEvaluación de Sora 2 ProPublishedSynchronized video and audio
GPT‑Live 1Evaluación de conversación GPT‑LiveMain modelNaturalness and interruptions
GPT‑TranscribeEvaluación ASR GPT‑TranscribePublishedTranscription accuracy
text‑embedding‑3‑largeMTEB64,6 %Published average
03 / METHOD

How to read the directory

An organization, a product, and a model are not the same unit. Inferama separates them to avoid attributing capabilities or commercial terms to the wrong item.

01

Laboratory

Entity that develops or publishes the model and maintains its technical and safety documentation.

02

Model

Identifiable version with limits, modalities, and behavior that may change between releases.

03

Access channel

API, application, associated cloud, or commercial plan; each channel may have different pricing, retention, and limits.

04

Source and date

Every claim must retain the official page consulted and the verification date.

ANALYSIS

Related analysis

Sora 2 Pro in Video Benchmarks: What Blind Preference Measures—and What It Leaves Unanswered
ANALISIS

Sora 2 Pro in Video Benchmarks: What Blind Preference Measures—and What It Leaves Unanswered

An arena score reports the outcome of a comparison under a specific protocol; by itself, it does not prove that a model follows instructions better, maintains temporal continuity, or synchronizes audio and video. This guide separates what can be said about Sora 2 Pro from what the available benchmarks do not establish.

26 Sep 2026
GPT‑Live Conversation Evaluation: How to Measure Turn-Taking, Interruptions, and Task Resolution Without Mistaking Smooth Dialogue for a Reliable Agent
ANALISIS

GPT‑Live Conversation Evaluation: How to Measure Turn-Taking, Interruptions, and Task Resolution Without Mistaking Smooth Dialogue for a Reliable Agent

A guide to evaluating GPT‑Live‑1 in full-duplex voice conversations by separating turn dynamics, comprehension, task success, and the safety of delegated actions. It proposes a reproducible protocol using controlled audio, traces, temporal annotation, and blinded human review.

23 Sep 2026
GPT‑Transcribe ASR Evaluation: How to Measure Text, Speakers, and Timestamps Without Hiding Errors That Break Downstream Workflows
ANALISIS

GPT‑Transcribe ASR Evaluation: How to Measure Text, Speakers, and Timestamps Without Hiding Errors That Break Downstream Workflows

A useful automatic speech recognition evaluation cannot be reduced to a single accuracy percentage. This guide proposes a protocol for separately measuring text fidelity, critical entities, speaker attribution, timestamps, segmentation, and the validity of the output consumed by an application. The goal is to compare versions and configurations reproducibly, then decide with evidence when to promote, restrict, or block a deployment.

23 Sep 2026
Localized AI image editing: how to verify that the model changes only what was requested and preserves the rest of the asset
GUIA

Localized AI image editing: how to verify that the model changes only what was requested and preserves the rest of the asset

A generative edit that is useful in production is not validated merely because the requested change looks correct. This guide proposes treating every retouch as a preservation contract: define what may change, what must remain, how to measure it, and when to require human review.

23 Sep 2026
GPT-Transcribe: what a team should revalidate before replacing its transcription system
NOTICIA

GPT-Transcribe: what a team should revalidate before replacing its transcription system

GPT-Transcribe is listed as a transcription model available through the OpenAI API. Before replacing an existing ASR system, teams should validate not only the resulting text, but also the technical contract that supports search, summaries, alerts, quotations, reviews, and compliance records.

22 Sep 2026
BenchCAD: What Reconstructing an Executable Mechanical Part Demonstrates—and What It Does Not Prove About Industrial Design
ANALISIS

BenchCAD: What Reconstructing an Executable Mechanical Part Demonstrates—and What It Does Not Prove About Industrial Design

BenchCAD 1.0 measures whether a system can generate, modify, or interpret CadQuery programs whose geometry can be executed and compared with a reference. It is a useful signal for parametric reconstruction, but it does not replace validation of tolerances, function, manufacturing, or integration into an engineering workflow.

22 Sep 2026
04 / SOURCES

Traceability