DECISION / COMPARATOR

Models face to face.

A common reading of capabilities, access, context, price and verification date.

Choose two or three models
O
Best suited for
Complex reasoning, research, code, documents, and agents with tools.
Worth monitoring
Higher price per token; requires minimum permissions and review for sensitive actions.
A
Anthropic

Claude Opus 5

Best suited for
Agentic coding, broad refactoring, enterprise knowledge, and visual work.
Worth monitoring
Responses tend to be lengthy, and active thinking can increase output tokens.
Best suited for
Low-cost agents, long-running software, and broad multimodal analysis.
Worth monitoring
The price is introductory, and using grounding may incur per-query charges.
CriterionGPT‑6 AstraClaude Opus 5Gemini 3.8 Flash
Identity and validity
OrganizationOpenAIAnthropicGoogle
API IDgpt-6-astraclaude-opus-5gemini-3.8-flash
StateAvailableAssetStable GA
ReleaseSeptember 3, 2026July 24, 2026September 2, 2026
Knowledge cutoffApril 30, 2026May 2026Not published in the source consulted
ProfileGeneralist frontierKnowledge and codeFast multimodal
Capacity and modalities
EntranceText · ImageText · ImageText · Image · Video · Audio · PDF
ExitTextTextText
Context window1,050,000 tokens1,000,000 tokens1,048,576 entry tokens
maximum output128,000 tokens128,000 tokens65,536 output tokens
Reasoninglow, medium, high, xhigh, and max effortAdaptive thinking · configurable effortlow, medium, and high levels
Relative latencyModerate; depends on reasoning effortModerate · Fast mode available in the APIFlash profile focused on speed and efficiency
Tools and integration
Published toolsFeatures · Web search · File search · Code interpreter · Hosted shell · Computer use · MCPUse of tools · Web search · Code execution · Computer useFeatures · Web search · Google Maps · Code execution · Computer use (preview) · File search · URL context
Structured OutingsCompatibleCompatibleCompatible
Batch processingCompatible · 50% discountCompatible · 50% discountCompatible · 50% discount
Prompt cacheCompatibleCompatible · up to 90% savings on readsCompatible
Fine-tuningNot compatibleNot published in the source consultedNot published in the source consulted
Standard API price
Entrance price10.00 USD / 1 M tokens5.00 USD / 1 M tokens0.75 USD / 1M tokens
Cached input1.00 USD / 1 M tokens0.50 USD / 1 M tokens0.075 USD / 1 M tokens
Cache write or storage12.50 USD / 1 M tokens6.25 USD / 1 M tokens (5 min) · 10.00 USD (1 h)0.50 USD / 1 M tokens per hour of storage
Starting price50.00 USD / 1 M tokens25.00 USD / 1 M tokens3.75 USD / 1M tokens
Batch input5.00 USD / 1 M tokens2.50 USD / 1 M tokens0.375 USD / 1 M tokens
Batch output25.00 USD / 1 M tokens12.50 USD / 1 M tokens1.875 USD / 1 M tokens
Relevant conditionsRequests with more than 272,000 input tokens apply price multipliers; tools may add per-call charges.Thinking is billed as output. Fast mode costs twice the base rate.Introductory pricing until December 31, 2026; from 2027, standard rates double.
Availability
Access channelsOpenAI API · Responses, Chat Completions, and BatchClaude, Claude API, AWS, Google Cloud, and Microsoft FoundryGemini API · Google AI Studio · stable version
Verified platformsOpenAI APIClaude.ai · Claude Code · Claude API · Amazon Bedrock · Google Cloud · Microsoft FoundryGemini API · Google AI Studio
Verified
Primary sourceOfficial GPT‑6 Astra page ↗Official Claude model documentation ↗Official Gemini 3.8 Flash model profile ↗

i «Not published» indicates that the primary sources consulted do not provide that data. Inferama does not fill gaps with estimates.

02

Comparable cost example

A request with 1 million input tokens and 250,000 output tokens, without caching or tools.

OpenAI

GPT‑6 Astra

Standard
22,500 USD
Batch
11,250 USD
Anthropic

Claude Opus 5

Standard
11,250 USD
Batch
5,625 USD
Google

Gemini 3.8 Flash

Standard
1,688 USD
Batch
0,844 USD

i It is an arithmetic example, not a cost estimate per task. Reasoning, caching, tools, retries, and context tiers may change the bill.

03

How to make a decision with this table

The most capable model is not always the most suitable system for a specific workflow.

01

Filter by hard requirements

First rule out what does not support your modality, context, tool, region, or maximum budget.

02

Test your own tasks

Compare quality, latency, error rate, and recovery using real cases and a stable rubric.

03

Measure the entire system

Include tool calls, searches, caching, oversight, retries, and the cost of correcting failures.

04

Review validity

Prices and limits change. Check the date and open the primary source before signing up or migrating.

04

Sources consulted

Technical documentation, pricing, status, and transparency for each provider.