A
ANTHROPIC / claude-sonnet-5

Claude Sonnet 5

Coding, research, and professional tool-based workflows.

01 / SUMMARY

The essential

ORIENTATIONCost/capability balance

Primary use stated or inferred cautiously from official documentation.

CONTEXT1,000,000 tokens

Maximum input capacity when published by the source.

EXIT128,000 tokens

Documented maximum generation limit.

ACCESSClaude, Claude API and associated clouds

Verified availability channels.

MODEL TYPEText and reasoning

Primary functional family and verified specialties.

LICENSEProprietary

Declared terms for API access, model weights, or self-hosting.

02 / TECHNICAL SHEET

Limits and integration

API IDclaude-sonnet-5
Model typeText and reasoning
Access modelPaid proprietary
LicenseProprietary
DeploymentHosted API
Release30 JUN 2026
Knowledge cutoffJAN 2026
EntranceText · Image
ExitText
Context window1,000,000 tokens
maximum output128,000 tokens
ReasoningAdaptive thinking · configurable effort
Published toolsUse of tools · Web search · Code execution · Computer use
Structured OutingsSupported
Batch processingCompatible · 50% discount
Prompt cacheCompatible · up to 90% savings on reads
Fine-tuningNot published
Verified platformsClaude.ai · Claude API · Amazon Bedrock · Google Cloud · Microsoft Foundry
BEST SUITED FOR

Coding, research, and professional tool-based workflows.

WORTH MONITORING

Cost and performance depend on the reasoning level.

03 / CAPABILITIES

What it can do

01

Orientation

Coding, research, and professional tool-based workflows.

02

Context

1.000.000 tokens · 128.000 tokens

03

Tools and integration

Tool use · Web search · Code execution · Computer use

04

Access

Claude, Claude API and associated clouds

Input modalities
TextImage
04 / PRICES

Documented cost

ConceptWorthUnit/condition
Standard input2,00 USD / 1 M tokensStandard API rate
Cached input0,20 USD / 1 M tokensReading reused prefixes
Cache write or storage2,50–4,00 USD / 1 M tokensThe condition varies by provider
Standard output10.00 USD / 1 M tokensMay include reasoning tokens
Batch input1.00 USD / 1 M tokensAsynchronous processing
Batch output5.00 USD / 1 M tokensAsynchronous processing

Consult the primary source before budgeting for a deployment.

i Prices change and may depend on level, region or context length. Check the source before making a decision.

05 / EVALUATIONS

How to read the results

“

A public benchmark provides guidance, but does not replace an evaluation with your data, tools, budget, and error tolerance.

BenchmarkResultMetricSource
OSWorld-VerifiedCurve by effortSuccess and costView source ↗
BrowseCompCurve by effortAccuracy and costView source ↗

Inferama only highlights a “best result” when the metric, test set, configuration, and date allow for an equivalent comparison. The supplier's figures are presented as claims from its own source.

06 / VERSIONS

Chronology

Claude Sonnet 5

Version added to Inferama's verified catalog.

ANALYSIS

Related analysis

Claude Sonnet 5: how to measure the impact of its new tokenizer before migrating
ANALISIS

Claude Sonnet 5: how to measure the impact of its new tokenizer before migrating

Anthropic says the same text may produce more tokens with Sonnet 5 than with Sonnet 4.6. A test using your own corpus can show whether that reduces usable context or changes cost per task in a specific integration.

23 Sep 2026
BrowseComp: what an agent that finds a difficult fact on the web measures, and why getting it right does not prove it conducts reliable research
ANALISIS

BrowseComp: what an agent that finds a difficult fact on the web measures, and why getting it right does not prove it conducts reliable research

BrowseComp evaluates whether an agent can locate a brief, hard-to-find factual answer through persistent web browsing. It is a useful signal, but a limited one: a high score is not enough to establish research quality, source traceability, or reliability on open-ended tasks.

22 Sep 2026
Claude Sonnet 5 and Cohere Embed 4 in Multimodal RAG: How to Separate a Retrieval Failure from a Response Failure
COMPARATIVA

Claude Sonnet 5 and Cohere Embed 4 in Multimodal RAG: How to Separate a Retrieval Failure from a Response Failure

Claude Sonnet 5 and Cohere Embed 4 occupy different layers in a multimodal RAG system. This guide proposes a factorial experiment, with a frozen corpus and conditions, to measure separately whether the right evidence was retrieved and whether the response used it faithfully.

22 Sep 2026
Claude Fable 5.1: when caching and batch processing reduce cost per task—and when they only shift the bill
GUIA

Claude Fable 5.1: when caching and batch processing reduce cost per task—and when they only shift the bill

The price per million tokens is not enough to choose between a standard call, instruction caching, or batch processing. This guide provides a cost model per correctly completed task for Claude Fable 5.1, including formulas, scenarios, and minimum telemetry. The outcome depends on reusing context before it expires, controlling retries, and accepting—or not—the asynchronous timeline of Batch API.

22 Sep 2026
Anthropic: how to verify what changes when you use Claude through an API, partner cloud, or product
ANALISIS

Anthropic: how to verify what changes when you use Claude through an API, partner cloud, or product

Adopting Claude is not simply a matter of selecting a model family. The access channel determines which identifier is used, which retirement schedule governs it, which controls each party administers, and which documentation can support a technical decision. This guide offers a method for separating those layers without turning public policies or safety evaluations into guarantees they do not contain.

22 Sep 2026
07 / SOURCES

Traceability