A
ANTHROPIC / claude-haiku-4-5-20251001

Claude Haiku 4.5

High volume, subagents, and structured extraction.

01 / SUMMARY

The essential

ORIENTATIONHigh efficiency

Primary use stated or inferred cautiously from official documentation.

CONTEXT200.000 tokens

Maximum input capacity when published by the source.

EXIT64.000 tokens

Documented maximum generation limit.

ACCESSClaude, Claude API and associated clouds

Verified availability channels.

MODEL TYPEText and reasoning

Primary functional family and verified specialties.

LICENSEProprietary

Declared terms for API access, model weights, or self-hosting.

02 / TECHNICAL SHEET

Limits and integration

API IDclaude-haiku-4-5-20251001
Model typeText and reasoning
Access modelPaid proprietary
LicenseProprietary
DeploymentHosted API
Release15 OCT 2025
Knowledge cutoffFEB 2025
EntranceText · Image
ExitText
Context window200.000 tokens
maximum output64.000 tokens
ReasoningExtended thinking
Published toolsUse of tools · Web search · Code execution · Computer use
Structured OutingsSupported
Batch processingCompatible · 50% discount
Prompt cacheCompatible · up to 90% savings on reads
Fine-tuningNot published
Verified platformsClaude.ai · Claude API · Amazon Bedrock · Google Cloud · Microsoft Foundry
BEST SUITED FOR

High volume, subagents, and structured extraction.

WORTH MONITORING

Cost and performance depend on the reasoning level.

03 / CAPABILITIES

What it can do

01

Orientation

High volume, subagents, and structured extraction.

02

Context

200.000 tokens · 64.000 tokens

03

Tools and integration

Tool use · Web search · Code execution · Computer use

04

Access

Claude, Claude API and associated clouds

Input modalities
TextImage
04 / PRICES

Documented cost

ConceptWorthUnit/condition
Standard input1.00 USD / 1 M tokensStandard API rate
Cached input0,10 USD / 1 M tokensReading reused prefixes
Cache write or storage1,25–2,00 USD / 1 M tokensThe condition varies by provider
Standard output5.00 USD / 1 M tokensMay include reasoning tokens
Batch input0.50 USD / 1 M tokensAsynchronous processing
Batch output2.50 USD / 1 M tokensAsynchronous processing

Consult the primary source before budgeting for a deployment.

i Prices change and may depend on level, region or context length. Check the source before making a decision.

05 / EVALUATIONS

How to read the results

“

A public benchmark provides guidance, but does not replace an evaluation with your data, tools, budget, and error tolerance.

BenchmarkResultMetricSource
SWE-bench Verified73,3 %ResolvedView source ↗
Terminal-Bench41,75 %With reasoningView source ↗

Inferama only highlights a “best result” when the metric, test set, configuration, and date allow for an equivalent comparison. The supplier's figures are presented as claims from its own source.

06 / VERSIONS

Chronology

Claude Haiku 4.5

Version added to Inferama's verified catalog.

ANALYSIS

Related analysis

Claude Haiku 4.5: How to Calculate Cost per Usable Response
GUIA

Claude Haiku 4.5: How to Calculate Cost per Usable Response

The price per call does not, by itself, show how much an accepted response costs. This guide provides a reproducible formula for adding input and output tokens, retries, and review, with three illustrative scenarios for Claude Haiku 4.5.

23 Sep 2026
Claude Haiku 4.5: How to Operate a Low-Latency Lane After Haiku 3.5 Without Mistaking Minimum Retention for Guaranteed Continuity
ANALISIS

Claude Haiku 4.5: How to Operate a Low-Latency Lane After Haiku 3.5 Without Mistaking Minimum Retention for Guaranteed Continuity

Claude Haiku 4.5 can be a destination for fast workloads after Claude Haiku 3.5 was retired, but a safe migration is not solved by changing an identifier. Documented minimum availability through October 15, 2026 does not guarantee indefinite continuity. The operational criterion is to demonstrate, using your own traffic and corpus, queue latency, schema compliance, tool use, and a reversible replacement path.

23 Sep 2026
Cost in SWE-bench: how to calculate the price per resolved issue without hiding failures, retries, or evaluation
ANALISIS

Cost in SWE-bench: how to calculate the price per resolved issue without hiding failures, retries, or evaluation

A cost figure per resolved issue is useful only if it reveals everything that happened before the patch was obtained: failed attempts, token consumption, stopping rules, selection among runs, and evaluation resources. This analysis proposes a reproducible scorecard for reading and comparing SWE-bench results, with particular attention to SWE-bench Verified.

22 Sep 2026
Analyzing customer feedback with AI: how to turn thousands of comments into priorities without confusing frequency with impact
GUIA

Analyzing customer feedback with AI: how to turn thousands of comments into priorities without confusing frequency with impact

A path for product, support, and research teams that need to analyze reviews, tickets, surveys, or transcripts with AI without turning an automated summary into an unsupported decision.

22 Sep 2026
Claude Haiku 4.5 vs Claude Opus 5: when a cost premium can justify a better result
COMPARATIVA

Claude Haiku 4.5 vs Claude Opus 5: when a cost premium can justify a better result

The choice between Claude Haiku 4.5 and Claude Opus 5 should not be settled solely by per-token pricing or unrelated benchmarks. This comparison proposes a reproducible protocol for measuring accepted outputs, retries, human review, latency, and effective cost in structured classification, document synthesis, and complex technical review. The available sources can establish comparable conditions and technical constraints, but they cannot support empirical results without running and publishing the experiment.

22 Sep 2026
Anthropic: how to verify what changes when you use Claude through an API, partner cloud, or product
ANALISIS

Anthropic: how to verify what changes when you use Claude through an API, partner cloud, or product

Adopting Claude is not simply a matter of selecting a model family. The access channel determines which identifier is used, which retirement schedule governs it, which controls each party administers, and which documentation can support a technical decision. This guide offers a method for separating those layers without turning public policies or safety evaluations into guarantees they do not contain.

22 Sep 2026
07 / SOURCES

Traceability