Ilustración editorial para Claude Opus 4.5 en tareas largas: calcula el coste por trabajo terminado
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

The price per token is not the cost of completing a job

A price per million tokens lets you calculate part of the bill, but it does not answer the budgeting question that matters for a long workflow: how much does it cost to get an acceptable task result? To answer that, reconstruct every run, add the attempts that did not reach the expected result, and identify costs that are not part of Claude’s token bill.

It helps to distinguish three figures. The cost of a run is the amount associated with a call or a recorded unit of work. The cost of an accepted task includes all the runs needed for the result to meet the acceptance criteria. Total operating cost adds items such as external tools and human review, if the team chooses to include them. These are not interchangeable metrics: a task may require several calls, and a generated response may be discarded.

This guide provides a calculation method, not a universal spending forecast. Prices vary with configuration and can change. The goal is to use your own Claude Opus 4.5 data, set limits before running a workload, and revisit the figures whenever prices, usage patterns, or the workflow change.

02

Set a verifiable price before you do the math

Before multiplying tokens by prices, record which price you are using and the context in which it applies. Anthropic’s pricing documentation should be your reference for the current terms for Claude Opus 4.5. The model’s original announcement helps identify the price quoted at launch, but it does not replace checking current pricing: a historical figure is no guarantee that it still applies.

As a dated snapshot, Anthropic’s price list dated May 27, 2026, lists Claude Opus 4.5 on the Claude API at $5 per million input tokens and $25 per million output tokens under standard pricing. It also lists different rates for cache writes and reads and for batch processing. Those figures are examples tied to that price list, not a promise of future pricing or a recommendation to use a particular service mode.

The same list gives rates of $6.25 and $10 per million for cache writes, $0.50 for cache reads, and $2.50/$12.50 for batch input and output. Do not infer from those numbers which type of write applies to your case or which combinations are allowed: check the current definitions and terms in the pricing documentation. Region, access channel, and service mode can affect the scope of a price. If you use a provider other than the Claude API, verify that provider’s own terms.

Save the date and an internal reference to the source you checked in your spreadsheet or change log; do not just copy the number. When a price changes, keep the previous calculation with its date so you do not compare runs under different prices as if they were equivalent.

Details to record when setting a price

FieldWhat to record
Model and identifierClaude Opus 4.5 and the exact identifier used by the integration.
Channel and scopeClaude API or another channel; region or provider, where applicable.
Service modeStandard, cache, batch, or any other applicable service condition.
Prices and dateAmount per billing unit and the date it was verified.
TermsEligibility, combination, and duration rules that affect the price.
03

Reconstruct the bill by usage category

A useful estimate keeps separate the categories that may have different prices. The general formula for a run is to add, for each category, the recorded tokens multiplied by the applicable rate, then divide by one million if the rate is stated per million tokens. Do not round each line too early: keep the precision until you have the total.

For a standard run, calculate input and output separately. If there was cache use, record writes and reads in their corresponding categories rather than counting the same block as both standard input and cache. If batch processing applies to the request, use the current batch rates for the relevant categories. Do not assume every token type gets the same discount or that cache and batch can be combined.

Messages responses include usage information for the call, and the API reference documents fields for input, output, and cache-related usage. For audit purposes, retain the usage response or a faithful extraction of its values along with a task ID and attempt ID. If you need to reconcile spending over a period, the Messages usage report offers another way to check it. Make sure its filters and time window match those in your own records.

A per-run calculation does not replace invoice reconciliation. Differences can arise from the scope you are measuring, a price change, or a service mode other than the one you assumed. If your internal totals do not match the usage report, pause extrapolation and resolve the discrepancy first.

Per-run calculation template

Apply the verified rate to each row. The prices in the earlier example are included only to demonstrate the method and should not replace the rates currently applicable to your account.

CategoryTokensPrice per millionCalculation
Standard inputInput tokens billed as standardStandard input rateTokens × rate ÷ 1,000,000
Standard outputOutput tokens billed as standardStandard output rateTokens × rate ÷ 1,000,000
Cache writeTokens recorded as a writeRate applicable to that writeTokens × rate ÷ 1,000,000
Cache readTokens recorded as a readCurrent read rateTokens × rate ÷ 1,000,000
BatchEligible tokens in each categoryApplicable batch rateCalculate by category and add
04

Worked example: a multi-attempt analysis task

Suppose a workflow summarizes documents and prepares a response for internal review. One accepted attempt uses 40,000 standard input tokens, 80,000 cache-write tokens, 160,000 cache-read tokens, and 8,000 output tokens. To illustrate the calculation, use the figures from the dated price list cited above: $5 per million for standard input, $25 for output, $6.25 for the selected write, and $0.50 for reads. That write rate is selected only for this example; in production, you would need to confirm which rate applies to your specific configuration.

The accepted attempt comes to: 40,000 × 5 / 1,000,000 = $0.20 for input; 80,000 × 6.25 / 1,000,000 = $0.50 for the write; 160,000 × 0.50 / 1,000,000 = $0.08 for the read; and 8,000 × 25 / 1,000,000 = $0.20 for output. Token cost for that attempt totals $0.98.

Now add two earlier attempts associated with the same task. One failed attempt uses 20,000 standard input tokens and 4,000 output tokens: $0.10 for input plus $0.10 for output, or $0.20 total. A discarded result uses 10,000 input tokens and 2,000 output tokens: $0.05 plus $0.05, for a total of $0.10. The accepted task therefore required $1.28 in token costs, not $0.98. This illustrates why a successful run does not necessarily represent the cost of completed work.

If a person also spends six minutes reviewing the result and the team assigns an illustrative internal cost of $45 per hour to that time, review adds $4.50. If an external tool adds another $0.30, the total operating cost in this example is $6.08. These items are internal accounting assumptions; they are not part of Anthropic’s token rates. Replace them with your organization’s actual costs and state what you include.

Example summary

Illustrative amounts calculated using the dated price snapshot mentioned above. They are not an official budget or a guarantee of cost.

Attempt or itemToken costStatus
Accepted attempt$0.98Accepted
Failed attempt$0.20Included in the task
Discarded result$0.10Included in the task
Total token cost$1.28One accepted task
Illustrative human review$4.50Internal cost, not a model rate
Illustrative external tool$0.30External cost
Illustrative total operating cost$6.08With the listed items
05

Count attempts that did not produce an acceptable result

Clearly define what “accepted” means before measuring. It might mean passing human review, a structured validation, or an agreed business condition. If the criterion changes between teams or versions of the workflow, cost per accepted task is no longer comparable. Record the outcome with each attempt, not only at the end of the process.

Associate all calls with a persistent task ID and assign an attempt number. Include failed runs, discarded responses, regenerations, recovery calls, and canceled jobs that have already consumed billable usage. Excluding them makes the apparent task cost lower than the spending the team actually incurs.

For a group of tasks, add the costs associated with those tasks and divide by the number of accepted tasks. Also report how many tasks were not accepted or are still pending; otherwise, an average can hide the fact that many attempts never reach completion. When a task has no acceptance, retain its cost and status. Do not simply delete it because it does not contribute to the accepted-task denominator.

Minimum recording workflow

  1. 01Assign a unique ID to each task and another ID to each attempt.
  2. 02Record the model, configuration, service mode, and run date.
  3. 03Save the reported input, output, and cache tokens for every call.
  4. 04Record the status: accepted, failed, discarded, canceled, or pending.
  5. 05Link external costs and review time if they are part of the operating metric.
  6. 06Group runs by task and calculate the total without removing earlier attempts.
06

Cache and batches: options to validate against your workflow

Cache may matter when a workflow reuses context. To find out whether it reduces total cost, compare the bill for a real sequence with and without reuse, using the write and read tokens actually recorded. The existence of a different read rate does not by itself prove that a particular task saves money: the amount reused, required writes, context structure, and current terms all matter.

Batch processing is worth evaluating when the work can wait and meets the documented requirements. Do not choose it just because a price in a published list looks lower. Check the mode’s turnaround and operational rules, whether the workflow can tolerate the latency, and which feature combinations are allowed. Your comparison should include both cost and time to acceptance.

To compare options, run a representative set of tasks with your current configuration and the alternative you are considering. Keep instructions and acceptance criteria constant and, where possible, use similar inputs. Calculate cost per accepted task and the share of repeated attempts for each variant. If the outcome depends on different tasks or changing criteria, do not attribute the difference to cache or batch without further checking.

Decision table for service modes

Use this as guidance on what to measure; check current documentation for eligibility and rates.

OptionWhen to evaluate itWhat to check
StandardWhen you need a straightforward baseline or an immediate response.Input and output tokens, price, and access channel.
CacheWhen context is repeated between calls and may be reused.Write and read tokens, applicable price, and reuse terms.
Batch processingWhen the task can run asynchronously.Eligibility, terms, timing, per-category prices, and compatibility with the configuration.
07

Set maximum budgets and stop rules

A useful limit has at least two levels: a maximum per run and a maximum for the complete task. The first bounds a call that grows beyond expectations; the second limits cumulative spending across retries. Before setting them, estimate the cost of a representative task and observe variability in a pilot set. An average alone is a weak basis if a few tasks cost much more.

Define limits in terms the system can enforce or monitor. These might include a maximum output-token count per call, a maximum number of attempts, a cumulative estimated spend per task, or a stop condition when the budget is exceeded. Whether each limit can be enforced depends on the integration architecture; do not assume the price itself will stop a run.

Add alerts for aggregate spending over a period and review the tasks that contribute the most to cost. If a control relies on a real-time estimate, document how it is calculated and how it will later be reconciled with recorded usage. An alert is not a guarantee of a block: distinguish notifications from technical controls that actually interrupt or reject work.

When choosing a threshold, consider the value and criticality of the work, the cost of review or failure, and the amount of usage variation the team is willing to accept. There is no universal monetary limit for Opus 4.5. The right amount depends on the task, volume, and each organization’s budget rules.

Operational rules before launching a workload

  1. 01Set an approved limit per attempt and a cumulative limit per task.
  2. 02Set a maximum number of retries and define which conditions allow them.
  3. 03Decide what counts as accepted and who resolves ambiguous cases.
  4. 04Enable alerts for the period budget and assign an owner.
  5. 05Define the action at the limit: stop, send for review, or request approval.
  6. 06At the end, reconcile the internal calculation with available usage records.
08

Reusable template and checklist

A spreadsheet can serve as an initial control if it preserves both the detail and the total. Use one row per call and add task and attempt IDs. Record the date, model identifier, channel, service mode, token categories, applied prices, calculation, result status, and any operating items you choose to include. Keep provider rates separate from internal costs so a change does not mix unlike items.

For each row, calculate token cost by multiplying the tokens in each category by its rate and dividing by one million, then add the results. Next, sum the rows associated with a task. Mark acceptance according to a shared rule and calculate the average cost per accepted task for the period, alongside the number of pending or unaccepted tasks. Keep a copy of the rates and the verification date so you can reconstruct the result months later.

When investigating a cost increase, first determine whether prices, context length, output volume, cache frequency, retries, or the acceptance rate changed. That analysis distinguishes an increase caused by pricing from one caused by the workflow itself. If you change several conditions at once, record the change and avoid attributing causality to a single variable without a controlled comparison.

For information about prices and the Claude Opus 4.5 model page, consult the pricing and model routes within Inferama. Use Anthropic’s official references to verify rates and usage fields before publishing budgets or running a significant workload. A calculated figure is only as reliable as the usage data, acceptance criteria, and dated price behind it.

Open questions

  • Claude Opus 4.5 prices and terms can change; check Anthropic’s official documentation when closing the guide and before each budget estimate.
  • The price list dated May 27, 2026, is a dated reference and does not establish that those amounts remain current at the time of use.
  • The supplied source summary does not detail the exact terms associated with each cache-write price; verify which category applies to each configuration.
  • Eligibility, compatibility, turnaround, and possible differences by region or channel must be confirmed for the specific account and service mode.
  • The cost of tools, in-house infrastructure, and human review depends on the organization; the example figures are illustrative assumptions, not verified rates.
09

Keep exploring

09

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction