What changes—and what the token figures do not prove
Anthropic’s migration guide says Claude Sonnet 5 uses a new tokenizer. For the same input text, the token count may be approximately 30% higher than with Claude Sonnet 4.6, although the precise increase depends on the content. The model announcement describes the variation as an approximate range of 1.0 to 1.35 times, also dependent on the type of text. These are two indicative descriptions of the same change, not a promise that every prompt will grow by the same proportion.
The first consequence is accounting-related: text that previously produced a certain number of tokens may register more units with Sonnet 5. The second is operational: if an application uses nearly all of its context window, the same material could take up a larger share of the available budget. The migration guide also warns that a max_tokens value tuned for Sonnet 4.6 may truncate an equivalent response when used with Sonnet 5. So it is not enough to swap the model identifier and assume that earlier measurements still describe the new behavior.
The figure alone does not show that costs will rise by the same proportion, that every input will grow by 30%, that the model will perform worse, or that its announced maximum context will shrink. Cost depends on the applicable rates and the billed input and output tokens, as well as any caching treatment. The usefulness of a response needs to be evaluated separately. A higher token count describes a difference in representation, not the quality of the result.
The release notes say Sonnet 5 supports a one-million-token context window and a maximum output of 128,000 tokens. Even if the nominal context limit is higher than that of a previous configuration, how much text fits within that limit depends on how the content is tokenized. It is therefore important to distinguish the model’s limit from an application’s effective ability to fit documents, instructions, conversation history, and tools within that limit.
What to measure in a real integration
A useful comparison does more than place two input-token counts side by side. It looks at the full request budget: system instructions, the user’s question, retrieved documents, conversation history, tool definitions, and the generated response. For a task with extensive context, an increase in document tokens may matter more than one in a short question. In an application with long outputs, generation limits and the max_tokens value also deserve a specific check.
Token count is an intermediate metric. To find out whether the change affects the product, track at least the requests that exceed the planned budget, truncated responses, limit errors, retries, discarded responses, and accepted results. Define an accepted task before running the test—for example, a response that meets the team’s predefined validity, format, and quality criteria. Without a definition in advance, it is easy to report an execution as cheaper when it simply failed to complete the work.
Cost per accepted task can be expressed as the total spend for the included executions divided by the number of tasks that meet the acceptance criterion. The numerator should include failed responses, retries, and discarded generations that incurred billable usage. Also specify how caching is counted, which rate was applied, and which access channel was used. This is not an official rate; it is an operational measure proposed for comparing an integration.
Repeated prompts should also be distinguished from new ones. If part of the context is reused through caching, the resulting charge may differ from that for an input processed without that treatment. Anthropic’s official pricing pages describe rates and cache multipliers, but a team needs to check the values currently in effect for its model, region, and platform when it runs the evaluation. Do not carry a historical figure, or a figure from another platform, into a current forecast without checking it.
Build a corpus that lets you attribute the change
The evaluation corpus should represent actual usage and remain frozen throughout the comparison. A stratified sample could separate English, other languages used by the product, code, long documents, short text, and repeated content. Stratifying the corpus does not assume those categories will always change in the same direction. It is specifically intended to reveal whether the effect varies between them. If the product handles conversations, include representative histories. If it retrieves documents, preserve the queries and the passages it actually sends to the model.
Save the exact text of every input, its composition, the count obtained with each model, and the task result. Avoid normalizing whitespace, changing instructions, or updating documents between runs unless that transformation is explicitly part of the test. Record the configuration, the date, and the access channel as well. That way, if token counts change, you can investigate the difference rather than vaguely attributing it to the migration.
Quality comparisons need identical criteria and consistent review. Where possible, evaluate results without telling the reviewer which model produced them. The goal is not to declare one version generally superior, but to check whether the replacement preserves the outcomes this team needs for its tasks. A higher token count can coexist with similar or better quality; a lower cost can coexist with more failures. That is why consumption and acceptance need to be considered together.
A minimum test-corpus matrix
Adapt the categories to your actual request mix. Use the same inputs for both versions.
| Stratum | What to include | What to observe |
|---|---|---|
| Languages | English and the other languages used in production | Input tokens, failures, and acceptance by language |
| Code | Representative snippets and application tasks | Token count, output format, and technical validity |
| Length | Short and medium requests, plus long documents | Context headroom, truncation, and limit errors |
| Repetition | New inputs and reused content | Billed tokens and the effect of caching treatment |
A protocol for comparing Sonnet 5 with Sonnet 4.6
Before running the test, decide what question you need to answer. It might be whether current prompts still fit with enough headroom, whether output limits need to be retuned, or whether spend per accepted task changes materially. Do not combine these into a single conclusion: each question requires a different metric. Also define what would count as material for the product, such as a reduction in context headroom that forces you to remove necessary content, or a cost change that exceeds the team’s internal threshold.
Run the same inputs against both versions and keep instructions, tools, generation parameters, and acceptance rules equivalent as far as the access channel allows. If a feature has no equivalent between versions, document the difference and do not automatically attribute its effect to the tokenizer. Record input and output counts, completion statuses, errors, and cache use separately. For each run, save the cost calculated using the applicable rate for that channel.
Then compare results by stratum, not just through an overall average. An average can conceal a small increase for English alongside a larger increase for some long documents, or a problem that appears only in requests close to the limit. Review the outliers and failed tasks. If the acceptance rate changes, inspect individual examples to distinguish content errors from cutoff at a limit, invalid formats, or configuration changes.
Repeat the test when a relevant variable changes. If rates, documentation, the corpus, or the caching system are updated, an earlier figure no longer describes the new conditions exactly. Keeping a dated version of the test set and calculation makes it possible to compare results over time without confusing model changes with changes in the surrounding environment.
A reproducible evaluation sequence
This protocol compares the same work under controlled conditions and keeps consumption metrics separate from outcome metrics.
- 01Freeze a representative corpus and divide it by language, content type, length, and repetition.
- 02Define in advance what counts as an accepted task, the quality criteria, and internal context and cost thresholds.
- 03Run the same inputs with Sonnet 4.6 and Sonnet 5 using equivalent instructions and configurations; record any unavoidable differences.
- 04Save input and output counts, errors, truncations, retries, cache use, acceptance results, and the rate applied.
- 05Compare by stratum and calculate total cost divided by accepted tasks, including discarded and failed executions that incurred usage.
- 06Repeat and date the evaluation when the corpus, rates, configuration, or access channel changes.
Separate tokenization, rates, and access channel
To attribute a change to the tokenizer, comparing an invoice from before migration with one from after is not enough. Pricing may have changed, cache use may differ, the prompt may have been edited, and responses may be a different length. The clearest comparison keeps three questions separate: how many tokens the same content registers, how much is billed under current conditions, and how many accepted tasks each configuration produces. A difference in one of these measures does not automatically explain the others.
Do not assume that every access channel offers identical conditions. Amazon Bedrock’s documentation provides service-specific information about the model’s availability, endpoints, regions, APIs, and billing. Anthropic’s platform documentation is the relevant reference for direct API access. The verified sources available here do not establish that every detail of counting, caching, limits, or billing is identical across channels. If your application uses a third party, measure it there and consult that provider’s current documentation and terms rather than extrapolating from results on the direct API.
Apply the same caution to prices quoted in announcements. An announced price or historical comparison does not necessarily describe what is billed today on every platform and in every region. Use current pricing documentation to convert measured usage into cost, and keep a record of the date and channel for the rate you used. If you cannot keep conditions equivalent, describe the comparison as an evaluation of the complete system, not as an isolated measurement of the tokenizer.
How to interpret different results
An observed difference points to the next check; it does not identify a cause by itself.
| Result | Priority check | Conclusion not to jump to |
|---|---|---|
| More tokens, similar cost | Rates, caching, output length, and channel | That the tokenizer has no operational effect |
| More tokens and lower acceptance | Truncation, limits, and configuration changes | That the token increase alone caused the decline |
| Higher cost per task | Discarded executions, retries, and current rates | That the per-token rate increased |
| Different results across channels | Accounting, limits, and platform-specific documented conditions | That the difference comes from the model or generalizes everywhere |
Decide whether to migrate, adjust, or temporarily keep the previous version
A migration is better supported when testing shows that required tasks continue to meet acceptance criteria, prompts fit with sufficient headroom, and cost per accepted task stays within the team’s defined limit. If counts rise but do not cause truncation or a material change in the budget, the increase may be an accounting variation that does not require an architectural redesign. Even so, update forecasts and dashboards that depended on earlier counts.
If the problem appears only in requests close to the limit, it may be enough to review those cases: remove redundant context, adjust document selection, or retune max_tokens when the application needs longer responses. Validate these changes against the same test tasks, because changing prompts or content introduces a new variable. Do not remove important context solely to compensate for a higher count; the result still has to be acceptable.
If the comparison reveals frequent failures, higher cost per accepted task, or inconclusive results because of channel differences, a reasonable response may be to postpone migration for that integration, split the migration by task type, or roll it out in a controlled way. Temporarily retaining Sonnet 4.6 does not mean Sonnet 5 is generally worse. It means the team’s evidence does not yet support changing this particular use case. The decision should also take account of the availability and current conditions of both versions.
Document both the decision and the conditions that could make it invalid. A conclusion such as “migrate” is useful only when accompanied by the version evaluated, corpus, channel, rate, and date. Without those details, the figures are no longer reproducible, and a later change in pricing or documentation can make a reasonable comparison outdated.
Limits of the conclusion
Anthropic’s published information provides a reference point for starting a test, not a substitute for each organization’s corpus. The approximate token range varies by content, so a sample dominated by short text may not represent long documents, code, English, or other production material. Nor can the range tell you the final cost of a task, because that depends on the mix of inputs and outputs, the rate applied, and each integration’s caching treatment.
Conditions can change over time. A model announcement, migration page, and pricing page may be updated on different dates. Before deployment, check the current official documentation and record which version and channel were evaluated. The sources consulted do not establish universal equivalence in counts, limits, and billing between Anthropic’s API and third-party services. Any conclusion about those details requires verification for the specific channel.
The useful question, then, is not whether Sonnet 5 always consumes a fixed percentage more tokens. It is how resource use changes for the workload you plan to migrate and whether that change affects the acceptable outcome. A controlled experiment can answer that with less risk than extrapolating from a general figure. To explore other options in the catalog, start with the model index; a migration evaluation should still use consistent conditions and criteria for each case.
Open questions
- The precise token-count change for a particular prompt set cannot be inferred from the published general range; it requires measuring the team’s own corpus.
- The source notes do not provide enough current numerical pricing values to calculate the cost of a specific integration.
- Universal equivalence in counting, caching, limits, or billing between Anthropic’s direct API and Amazon Bedrock or other channels cannot be inferred.
- Rates and documentation may change; operational conclusions should be dated and reviewed before deployment.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction