Primary use stated or inferred cautiously from official documentation.
GPT‑5.6 Sol
Complex work, coding, and tool-using agents.
The essential
Maximum input capacity when published by the source.
Documented maximum generation limit.
Verified availability channels.
Primary functional family and verified specialties.
Declared terms for API access, model weights, or self-hosting.
Limits and integration
| API ID | gpt-5.6-sol |
|---|---|
| Model type | Text and reasoning |
| Access model | Paid proprietary |
| License | Proprietary |
| Deployment | Hosted API |
| Release | JUL 2026 |
| Knowledge cutoff | 16 FEB 2026 |
| Entrance | Text · Image |
| Exit | Text |
| Context window | 1,050,000 tokens |
| maximum output | 128,000 tokens |
| Reasoning | none, low, medium, high, xhigh, and max effort |
| Published tools | Features · Web search · File search · Computer use |
| Structured Outings | Supported |
| Batch processing | Compatible · 50% discount |
| Prompt cache | Compatible · 90% discount on reads |
| Fine-tuning | Not published |
| Verified platforms | OpenAI API |
Complex work, coding, and tool-using agents.
Cost and performance depend on the reasoning level.
What it can do
Orientation
Complex work, coding, and tool-using agents.
Context
1.050.000 tokens · 128.000 tokens
Tools and integration
Features · Web search · File search · Computer use
Access
OpenAI API · Responses and Chat Completions
Documented cost
| Concept | Worth | Unit/condition |
|---|---|---|
| Standard input | 4,00 USD / 1 M tokens | Standard API rate |
| Cached input | 0,40 USD / 1 M tokens | Reading reused prefixes |
| Cache write or storage | 5.00 USD / 1 M tokens | The condition varies by provider |
| Standard output | 20,00 USD / 1 M tokens | May include reasoning tokens |
| Batch input | 2,00 USD / 1 M tokens | Asynchronous processing |
| Batch output | 10.00 USD / 1 M tokens | Asynchronous processing |
Consult the primary source before budgeting for a deployment.
i Prices change and may depend on level, region or context length. Check the source before making a decision.
How to read the results
A public benchmark provides guidance, but does not replace an evaluation with your data, tools, budget, and error tolerance.
| Benchmark | Result | Metric | Source |
|---|---|---|---|
| Agents' Last Exam | 52,7 % | Accuracy | View source ↗ |
| SWE-Bench Pro | 64,6 % | Resolved | View source ↗ |
| Terminal-Bench 2.1 | 88,8 % | Accuracy | View source ↗ |
Inferama only highlights a “best result” when the metric, test set, configuration, and date allow for an equivalent comparison. The supplier's figures are presented as claims from its own source.
Chronology
GPT‑5.6 Sol
Version added to Inferama's verified catalog.
Related analysis
AutomationBench: What an Agent That Leaves a Business Workflow in the Right State Demonstrates—and What It Does Not
AutomationBench evaluates whether an agent can complete workflows across simulated SaaS applications and bring them to a verifiable final state. This guide explains what that evidence means, how to read its metrics, and what additional testing is needed before extrapolating a result to a real company.
22 Sep 2026 ↗ PAPERAgents’ Last Exam: what it measures about a real-work agent—and why its pass rate does not mean “automating a job”
Agents’ Last Exam evaluates agents on professional tasks within operating-system environments and with verifiable outcomes. Its design is useful for studying workflow execution, but an aggregate success figure does not demonstrate that a job can be automated. Interpreting it requires details about the tasks, environment, harness, model, tools, budget, and evaluated version.
22 Sep 2026 ↗