Ilustración editorial para Claude Haiku 4.5: cómo calcular el coste por respuesta utilizable
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

What cost are you calculating?

The price of a call, the spend on a task, and the cost of an accepted response are different measures. To budget for a real workload using Claude Haiku 4.5, start by defining the denominator: do you want to know the cost of each attempt sent to the model, each response received, each result that passes your acceptance rules, or each task completed after human review?

The distinction matters when a response may be discarded, corrected, or requested again. A call that returns text consumes tokens even if that text is not used. If a team divides total spend only by successful calls, it can hide failed attempts. If it divides by all calls, it does not show how much it costs to produce an output that is actually useful to the process.

In this guide, an “accepted response” means an output that passes the criteria the team has defined for the task. It might be a valid label, an object that conforms to a schema, or a summary that passes review. It does not mean the model guarantees accuracy, that the output is safe for every use, or that no further approval is needed. It is an operational unit for calculation, not a quality guarantee.

The published rate is only one component of the budget. The calculation can list inference, retries, external validation, human review, and other system resources separately. Keeping these items distinct makes it easier to compare runs and explain where a figure comes from, without presenting an estimate as a guaranteed invoice.

02

Claude Haiku 4.5 pricing and its scope

Anthropic lists standard Claude Platform pricing for Claude Haiku 4.5 at USD 1 per million input tokens and USD 5 per million output tokens. This is a useful basis for a reproducible estimate, but record it alongside the channel used, the exact model, and the date you checked it. Do not assume that pricing through a third-party channel or another service arrangement is identical.

The provider’s documentation also describes different pricing options, including those associated with batch processing and prompt caching. The calculations in this guide do not include them: the aim is to explain the calculation using the stated standard rates, not to determine which option applies to a particular workload. Before budgeting, check the current price list and the terms that apply to the channel you plan to use.

The model identifier matters both for reproducing a test and for confirming that runs correspond to the same version. The lifecycle documentation checked for this guide identifies the model as claude-haiku-4-5-20251001 and shows it as active, with no retirement date announced at the time of that check. That status can change; it should not be treated as a guarantee of future availability.

When verifying rates, record which pricing you applied and preserve the date of the check. If you intend to use the budget over several months, schedule another review of prices and terms before turning the projection into a commitment. A change in the unit price affects the result even if the number of tasks and the observed token counts remain the same.

Details to record alongside a rate

Record each detail in the same worksheet or report as the calculation. This distinguishes a published reference from your own measurements.

DetailWhat to recordWhy it matters
ModelExact identifier used in testingMakes it possible to reconstruct which model generated the outputs.
ChannelDirect API or another access channelAvoids assuming that every channel has the same rate.
RateUnit price for input and outputThese are different prices applied to different kinds of usage.
DateDate prices and terms were checkedA previously checked figure can become outdated.
Pricing optionStandard or another applicable optionPrevents mixing rates that have different terms.
03

The formula: inference cost per accepted response

Using the stated standard prices, estimate the inference cost of a text attempt as follows: (input tokens × USD 1 / 1,000,000) + (output tokens × USD 5 / 1,000,000). The formula treats the two sides of the interaction separately because each million tokens has a different price. For multiple calls, add the estimated cost of every attempt, not just the attempts that ended in acceptance.

Then divide total inference spend by the number of accepted responses. For a workload with different task types, first calculate cost and acceptance by task type, then combine the results weighted by volume. A simple average of percentages can distort the cost when tasks differ substantially in token use or frequency.

When aggregate data is available, a useful approximation is to multiply the average cost per attempt by the average number of attempts required to obtain an accepted response. Use this approximation only if you explain how you derived the average and confirm that it represents the workload being analyzed. If the first attempt and retries have different lengths, calculate the spend for each group separately rather than assuming every attempt costs the same.

Observed token counts should come from representative runs, not from an idealized length chosen to make the projection look favorable. The Messages API reference documents usage fields for input and output; these data can be used to compare recorded usage with the calculation assumptions. Preserve the unit and measurement period so tokens are not confused with characters, words, or requests.

A reproducible procedure

Apply the same criteria to every scenario and keep the source data.

  1. 01Define what counts as an accepted result and what conditions trigger a retry.
  2. 02Fix the model, channel, and rate to be used in the calculation.
  3. 03Measure input and output tokens per attempt in a representative sample; separate first attempts from retries if their usage differs.
  4. 04Calculate the inference spend for each attempt using the applicable rate, then add the costs of all attempts.
  5. 05Divide total spend by accepted responses in the same period, and list human or infrastructure costs in separate categories.
04

Three illustrative scenarios

The examples below show how the order of magnitude changes as token volume and the share of accepted attempts vary. They are hypothetical calculations, not published measurements of Claude Haiku 4.5 or performance predictions. They assume that attempts have the same average length, that each attempt has a constant probability of acceptance, and that retries continue until an attempt is accepted. Under those assumptions, expected inference cost per accepted output is the average cost per attempt divided by the acceptance rate.

For a brief classification task, suppose each attempt uses 600 input tokens and 80 output tokens, with a 95% acceptance rate. At the stated rates, an attempt costs USD 0.001: USD 0.0006 for input and USD 0.0004 for output. Under the assumptions above, expected inference cost is approximately USD 0.00105 per accepted response.

For structured extraction, suppose each attempt uses 1,800 input tokens and 450 output tokens, with an 85% acceptance rate. An attempt costs USD 0.00405: USD 0.0018 for input and USD 0.00225 for output. Expected cost per accepted response is approximately USD 0.00476. This example does not add a separate charge for schema validation; record validation as an external cost if it actually consumes billable resources.

For a bounded synthesis task, suppose each attempt uses 3,000 input tokens and 1,200 output tokens, with an 80% acceptance rate. The cost per attempt would be USD 0.009, and expected cost per accepted response would be approximately USD 0.01125. At the assumed rates, output accounts for two-thirds of the attempt cost. The example illustrates why counting requests alone is not enough: response length also changes the budget.

Hypothetical scenarios at standard rates

Amounts in USD per attempt and per accepted response. Human review, separately billed validation, infrastructure, and other services are excluded.

TaskInput tokensOutput tokensAssumed acceptanceCost per attemptEstimated cost per accepted response
Brief classification6008095%0.001000.00105
Structured extraction1,80045085%0.004050.00476
Bounded synthesis3,0001,20080%0.009000.01125
05

Which variables drive the result?

In these examples, each output token has a unit price five times that of each input token. That does not mean output always accounts for the largest share of a bill: it depends on how many tokens there are of each type. In a request with extensive input and a very short response, input can still be a significant part of the cost. In a task that produces lengthy text, output may dominate.

Acceptance affects cost per result even when it does not change the price of an attempt. In the simplified model, a lower acceptance rate means more expected attempts for each acceptance. This relationship does not describe every system universally: there may be retry limits, alternative routes, manual review, or different reasons for rejection. For that reason, when planning a real workload, calculate from recorded attempts and accepted outcomes in the same sample wherever possible.

Output length can also vary by request type and workload behavior. A maximum output limit is not a forecast of average usage. For budgeting, measure observed tokens and monitor whichever percentile matters to you. An average may be insufficient if long outputs are frequent or costly for the downstream process.

Do not attribute every failure to the model without classifying it. A rejection may result from a strict schema, incomplete input, a service interruption, or a business rule. Separating these causes helps identify what to fix and avoids inflating the model retry rate with validation or integration problems that have other solutions.

What to measure first

Prioritize your own measurements before optimizing a variable based on intuition.

Observed signalWhat to investigateWhat not to conclude automatically
Many long responsesOutput-token distribution by task typeThat every request needs a shorter response.
Many discarded attemptsRejection causes and the usage of each retryThat every discard is caused by the model.
Extensive inputTokens sent and the content the model needsThat context can be removed without affecting the task.
Difference between spend and total costValidation, review, tools, and infrastructureThat cost per token explains the whole process.
06

Retries, validation, and human review

Inference cost for an accepted response should include every attempt that consumed tokens before the result was accepted. If the first attempt is discarded and another is requested, both contribute to spend. If the policy sets a maximum number of retries, measure final acceptance and accumulated tokens under that specific policy; a formula that assumes unlimited retries will not describe that operation.

External validation can carry a cost even when it does not call the model again. A local format check can consume compute resources; validation through another tool can incur charges; manual intervention consumes staff time. Do not mix these costs with Claude Haiku 4.5 token usage. Record them separately and, if there is an internal hourly rate, state how human time was converted into money.

Human review may apply to every response, only a sample, or only uncertain cases. In each case, state what proportion was reviewed, how long the review took, and what share was ultimately accepted, corrected, or discarded. Without those data, review cost cannot be inferred from the API rate.

For many teams, the most useful measure has two views: inference cost per accepted response and total process cost per completed task. The first helps explain model usage. The second includes the components required to deliver the result in the organization’s actual context. Both should use the same period, volume, and definition of success.

07

A template for monthly projections

To project monthly spend, separate tasks by type and estimate monthly volume using usage data or an explicit assumption. For each type, record observed average input and output tokens, acceptance rate, attempts per result, and the share that requires review. If the task mix changes during the month, use a distribution by task type rather than one unweighted global average.

Multiply the expected monthly attempts for each type by its average inference cost per attempt, then add the task types together. Next, divide aggregate spend by the expected number of accepted responses to report average cost per usable output. To calculate total monthly process cost, add separately measured validation, review, tools, and infrastructure spend.

As a check, compare the projection with a sample of real records. Confirm that the token period matches the period used for accepted responses and that incomplete tasks have not been counted as successes. If projected and observed costs diverge, investigate changes in task mix, output length, retry rate, or applied pricing before changing the assumptions.

Monthly data-capture template

Complete one row per task type. Identify fields without measurements as assumptions; do not present them as observed data.

FieldRecord for each task type
Task type and volumeClassification, extraction, or synthesis; tasks expected during the month
Tokens per attemptObserved input and output averages, recorded separately
OutcomesTotal attempts, final acceptances, and acceptance rate
RetriesAverage number of retries and tokens consumed per retry
Applied rateInput and output prices, channel, pricing option, and verification date
External costsValidation, human review, tools, and infrastructure, recorded separately
Budget resultInference spend, cost per accepted response, and total process cost
08

Limits of the estimate and checks before budgeting

The scenarios in this guide do not predict usage for a particular application. Their token counts and acceptance rates are invented for illustration; they are not official averages or comparative quality results. Replacing them requires a sample of your own that reflects real inputs, instructions, constraints, validation rules, and retry policies.

Inference cost should not be treated as the full cost of a task. The examples exclude human review, externally billed validation, tools, infrastructure, and pricing options other than the stated standard rates. These components depend on the design and channel of each implementation.

Before approving a budget, recheck the model identifier and status, the current rate, and the channel’s scope. Prices can change, and caching or batch-processing options should not be substituted for standard pricing without checking their terms. Save the reference and the date you checked so someone else can reconstruct the estimate.

In short, a price per million tokens helps value usage, but does not by itself determine the cost of an output the team can use. Input and output length, discarded attempts, and acceptance rate determine inference cost per result; review and the rest of the process complete the operating cost. A useful figure states its denominator, data, and exclusions.

Pre-budget checklist

Before presenting a figure as a budget, verify the following.

  1. 01Has “accepted response” been defined in an observable way?
  2. 02Are input and output tokens recorded separately for first attempts and retries?
  3. 03Does the acceptance rate come from a representative sample and reflect the intended retry policy?
  4. 04Does the rate apply to the model and channel that will be used, and is it dated?
  5. 05Have pricing options other than the standard rate been checked separately or explicitly excluded?
  6. 06Are validation, human review, tools, and infrastructure listed as separate costs or exclusions?
  7. 07Have all figures that do not yet come from your own measurements been labeled as assumptions?

Open questions

  • Prices and terms can change; verify them before publication or budgeting and confirm that they apply to the specific channel.
  • The active status and absence of an announced retirement date reflect the documentation check described in the source material; they do not guarantee future availability.
  • The token counts, acceptance rates, and retry volumes in the three scenarios are illustrative assumptions, not measurements from a real implementation.
  • Validation, human review, tools, and infrastructure costs depend on the system and cannot be calculated from token rates alone.
  • The simplified calculation assumes attempts have constant length and acceptance probability; workloads with retry limits, variable lengths, or alternative routes should be estimated from their own records.
09

Keep exploring

09

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction