Ilustración editorial para Claude Sonnet 4.5 con herramientas: calcula el coste de una tarea, no solo de una llamada
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

The cost of a call is not the cost of a task

An integration may show a single user request and still send Claude several requests before the work is finished. The model may propose using a tool; the application runs that operation and returns its result; then the model processes the context again and decides whether to answer or use another tool. To estimate the task’s cost, count every request to the model—not just the initial request or final response.

This guide provides a spreadsheet template for Claude Sonnet 4.5. It separates input tokens, output tokens, and any charges that may apply to server-hosted tools. It does not include caching or the Batch API. It also does not assume that every task has the same number of rounds: a simple search, a multi-step workflow, and a failed attempt can have different costs.

The calculation covers use of the Claude API in a bounded workflow where the application runs tools. Do not automatically apply it to Bedrock, Google Cloud, or another platform: identifiers, rates, billing units, or availability schedules may differ. For the pricing guide, see the internal route pricing.index; for the model, Claude Sonnet 4.5; and for the provider, Anthropic.

02

Set the model, channel, and identifier before estimating

The product name alone is not enough to reproduce a calculation. For the Claude API, the identifier documentation distinguishes the alias claude-sonnet-4-5 from the dated version claude-sonnet-4-5-20250929. The supplied lifecycle page identifies that version as active and says it will not be retired before September 29, 2026. Treat this as a useful reference, not a guarantee of future availability: check the status just before publishing or reusing the spreadsheet.

Record the channel as well. Anthropic’s pages describe equivalent identifiers for other providers, but do not assume the same price, schedule, or usage format applies to all of them. The rate entered in the spreadsheet must match the actual channel and the date of the run. The pricing documentation is the reference for checking current rates, possible regional differences, and terms for associated platforms.

The verification materials supplied do not include current input and output rates or the amount of any regional differences, so this guide does not present them as confirmed figures. The template keeps those rates as editable fields. Before using the result as a budget, replace them with the values shown in the applicable documentation and note the date you checked them.

Control data for the spreadsheet

Fill in these cells using information from the run and the current rate. Do not reuse the spreadsheet after a change to the model, channel, or rate without updating those fields.

FieldWhat to recordWhy it matters
Model and identifierThe alias or dated version actually sentMakes the result reproducible and helps detect version changes
ChannelClaude API or another platformPrevents applying one provider’s rates or terms to another
Rate dateThe date the prices were checkedShows when the calculation needs to be updated
Input and output pricesThe published rate per unit for that channelThese are the multipliers for token usage
Hosted toolsBillable operations and their billing unitMay generate charges separate from tokens
03

Count a complete round of the tool-use cycle

In client-side tool use, a model response may contain a request to use a tool. The application performs the action and sends a tool result in a subsequent message; the model then processes the continuation. This exchange may repeat. The message requesting the tool is not the same as a completed task.

For every request to the model, count the input actually sent in that request. It may include the system message, previous messages, tool definitions, the user’s message, and earlier results that the application has kept in context. If a new request includes earlier content again, that content is part of the input calculation for that request again. Do not count a tool result as model input until you actually send it to the model.

Output is also counted per request. In an intermediate round, it may consist of a tool instruction and arguments; in the final round, it may be a response for the user. Do not assume an intermediate output has zero cost. For Claude Sonnet 4.5, tool use may add system tokens associated with the selected tool. The number depends on the configuration; check the pricing documentation. Do not add a generic value without verifying which tool and mode are in use.

A server-hosted tool requires an additional distinction. Its operations may have their own charges, and some loops may run internally in the service. In that case, each internal operation does not necessarily correspond to a client-visible model call. Track model requests and tool operations separately, and bill each concept according to its documented unit.

Sequence to capture in the log

One client-side cycle can contain more than one request to the model. The count should follow what actually happened, not the number of turns visible in the interface.

  1. 01The application sends the model the available context and tool definitions.
  2. 02The model returns a response, which may request a tool.
  3. 03The application runs the tool and records its result, duration, and any errors.
  4. 04The application sends the result to the model along with whatever context it chooses to retain.
  5. 05The model responds or requests another tool; the cycle continues until the completion criterion is met.
04

A calculation template by round and by task

Set up the spreadsheet with one row for each request to the model and a separate table or section for tool charges. For each model row, record billable input tokens, billable output tokens, input price, and output price, all for the recorded channel and date. If a hosted tool is used, also record each operation that the tool bills for. This lets you total a task’s cost without mixing units.

The token-based estimate for each call is: (input tokens ÷ rate unit) × input price + (output tokens ÷ rate unit) × output price. The rate unit must be the one published by the provider; if the price is quoted per million tokens, the unit is one million. For a task, add the result for every model call and then add any applicable tool charges separately.

Do not theoretically deduplicate the input total. If a system message, tool schema, or earlier result appears in three billable requests, count the usage recorded for each request, even if the text is identical. Conversely, do not add estimated tokens for content that was never sent. If reliable measurements are unavailable, mark the figure as an estimate and check it against a token-counting tool or the usage data returned by the API.

Suggested columns for a reproducible spreadsheet

Use one row per model call. Keep hosted tools in a separate table if they have different billing units.

Task and attemptRoundInputOutputInput priceOutput priceToken cost
Stable identifier1, 2, 3…Measured tokensMeasured tokensCurrent rateCurrent rate(input/rate unit × price) + (output/rate unit × price)
ToolOperationRecorded resultError or successBillable unitApplicable priceOperations × price
05

Three illustrative scenarios

The scenarios below explain how to structure the count; they do not predict universal usage. The token volumes are example numbers for practicing with the spreadsheet, not measurements of Claude or published rates. To calculate a real monetary cost, replace the illustrative token counts with your application’s logs and the prices verified for your channel.

Scenario A: a question answered without tools. Count the single request to the model, including the messages and context sent, and the generated output. If the application includes a tool definition even when it is not used, inspect the actual request and do not assume the definition was free or absent. There is one model round in this case, but that does not mean every query in the product has the same cost.

Scenario B: a question that needs a search. Count the first model call, the response requesting the tool, the search operation, and the subsequent call that includes the result. If a server-hosted tool is used, record its charge in the unit specified in its documentation; do not convert it into tokens. If the application integrates an external search on the client side, record the relevant external expense separately from Claude’s token bill.

Scenario C: an action with multiple sequential calls or a failed attempt. If the first tool returns an error and the application asks the model to correct it, add that call and any subsequent calls. If the process restarts as a new attempt, retain the attempt identifier and add its usage to the cost of the completed task. Reporting only the successful attempt would hide the spend required to reach the result.

06

Which variables can change the result

The number of rounds is often a decisive variable. A task that finishes after the first response may involve fewer calls than one requiring three model decisions. But do not attribute every difference to the model: it may depend on application logic, tool state, data quality, retry limits, or the success criterion.

The amount of context retained by the client matters too. Sending long definitions, previous messages, and large results can increase input in later rounds. Trimming the history may change usage, but it is only appropriate if the information needed for the task is preserved. Measure the effect in a controlled run instead of assuming context is summarized or billed only once.

Tool schemas and system tokens associated with tools add another variable. The precise amount depends on the tool and configuration, so check the pricing documentation and the request’s token count. A long tool result may also increase input for the next call, provided it is sent to the model.

Finally, distinguish failed tasks, automatic retries, and corrections requested by the user. A budget based on the average of completed tasks may underestimate spend if it omits failed attempts. Consider reporting both cost per attempt and aggregate cost per completed task, along with the measurement period and set of runs.

Diagnosing an increase in cost per task

Compare equivalent logs and change one variable at a time. This table can guide an investigation; it does not establish causality without measurements.

Observed signalWhat to checkUseful measurement
More model callsRound limits, retries, and completion conditionsRequests per attempted and completed task
Growing input in later roundsResent history, tool results, and schemasInput tokens per call and component
Growing outputArgument length, intermediate explanations, and final responseOutput tokens per round
Increase with no apparent token changeTool charges, channel, region, or platformBillable operations and applied rate
07

Check the spreadsheet against production and state exclusions

For an audit, retain a task and attempt identifier, model identifier, channel, request number, available tools, results sent to the model, and usage data for each response. Also record failed calls and operations that did not produce a useful response. Avoid storing personal data or secrets that are not needed for the analysis; cost metrics can be associated with pseudonymous internal identifiers.

The API documentation includes a token-counting endpoint that accepts messages, system instructions, and tools. It can help estimate a request before sending it. It does not replace checking the usage returned at runtime or the final bill, especially when there are multiple rounds, errors, or tool charges. Use the pre-send estimate for planning and actual logs for auditing.

Compare the calculated cost with the bill for an equivalent period and channel. Investigate discrepancies by looking for missing calls, misinterpreted rate units, regional differences, or separate tool charges. Keep a column for explicit exclusions. Caching and the Batch API are excluded from this calculation; they may require their own treatment and should not be added to the base workflow formula as if they were part of it.

A useful figure is not necessarily an exact prediction. Publish the method, pricing date, task set, number of attempts, and exclusions. If the sample covers little variety, present it as the result for that sample, not as the typical cost of any task using Sonnet 4.5.

Open questions

  • The supplied sources do not include current Claude Sonnet 4.5 input and output rates; verify them in the pricing documentation for the applicable channel and date.
  • No specific figures for regional differences or associated-platform pricing are supplied, so none are quantified.
  • System tokens associated with tool use may vary by tool and configuration; check current documentation and measure the specific case.
  • No specific hosted-tool rate or reproducible run with a bill is supplied, so the arithmetic example is not claimed to match a real charge.
  • The model’s status and retirement schedule should be checked again before publication; availability dates may differ between the Claude API and associated platforms.
08

Keep exploring

08

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction