Scope: compare modes without transferring prices between channels
Claude Fable 5.1 is offered through several channels, including Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS. This guide limits its comparison to the rules and prices documented for Anthropic’s direct platform. It is not valid to copy those rates to Amazon Bedrock or another cloud provider: the pricing documentation indicates that Bedrock and Google Cloud apply independent regional prices.
The model documentation places the launch of Claude Fable 5.1 on September 1, 2026, states a context window of one million tokens, and specifies a maximum output of 128,000 tokens. These limits can constrain both request design and the bill: context that fits in theory cannot necessarily be sent at the frequency, concurrency, or deadline the workload requires.
The objective, therefore, is not to establish that one mode is universally cheaper. It is to build a comparable calculation for a specific workload and measure the cost of a task that completes correctly. That cost includes tokens, retries, validation, failure recovery, and, when the product requires it, the operational cost of waiting for an asynchronous response. A lower token bill alone does not prove lower total cost, better results, or less human review.
What is billed in Claude Fable 5.1
On the direct platform, the Claude Fable 5.1-specific table documents USD 10 per million standard input tokens and USD 50 per million output tokens. A five-minute cache write costs USD 12.50 per million tokens; a one-hour write costs USD 20 per million. Cache reads cost USD 0.25 per million tokens. These figures come from the supplied documentation and must be verified again before a budget is implemented, because pricing and commercial terms may change.
A cache write is the initial processing of the prompt prefix that becomes available for reuse. A read occurs when a subsequent request matches that stored prefix and can use it. The time-to-live begins when the cache is created. The five-minute and one-hour durations are not capacity reservations, nor do they by themselves guarantee reuse: if the next call arrives too late, the relevant prefix changes, or the request does not obtain a hit, the expected savings do not materialize.
Output does not receive a cache discount as a result. An agent that produces lengthy explanations, patches, or reports can retain a high bill even with an excellent hit rate. The dynamic part of the prompt must also be counted: only the prefix that is actually reused can benefit from a cache read; instructions or context added afterward continue to be billed as standard input.
Batch API processes requests asynchronously, and the documentation states a 50% discount relative to standard prices. Caching and Batch discounts can stack. A price discount, however, does not make Batch a substitute for an interactive path: batch processing has an expiration window and its own operational discipline. The team must verify whether the work can wait and how much it costs to resubmit, correct, or investigate items that do not complete as expected.
Direct cost components to model
| Component | Documented rate per MTok | Control question |
|---|---|---|
| Standard input | USD 10 | What portion of the prompt is not reused? |
| Output | USD 50 | Does response length dominate the bill? |
| Cache write, 5 minutes | USD 12.50 | Will reuse happen before expiration? |
| Cache write, 1 hour | USD 20 | Does the longer lifetime prevent new writes? |
| Cache read | USD 0.25 | Does the prefix match and is a hit recorded? |
| Batch API | Documented 50% discount | Does the asynchronous timeline meet the operational requirement? |
The useful unit: cost per correctly completed task
Cost per call is an incomplete signal. A task may require several turns, an automated check, validation of a structured result, a recovery call, or a retry caused by a capacity limit. In addition, a response received after the product can no longer use it may have little operational value even if it was inexpensive.
Define a correctly completed task before comparing modes. For example, for an engineering agent it may be a proposal that passes tests and policy validation; for a document assistant, an answer that follows the required format and has retrievable evidence; for a nightly queue, a record processed before its delivery time and accepted by downstream controls. Record technical errors, invalid results, and work requiring human intervention separately.
For an analysis window, a practical metric is total cost of requests, retries, validation, and recovery divided by accepted tasks. If the goal is to isolate model cost, leave out salaries and external systems, but do not hide retries or corrective calls. If the goal is to choose a product architecture, add those components and the cost of missing a deadline through an explicit, reviewable assumption.
Measurement process per task
- 01Assign a task identifier that persists across the initial attempt, retries, and validation.
- 02Store, for every request, the standard input tokens, cache writes, cache reads, and output returned in API usage.
- 03Classify the outcome as accepted, retried, discarded, pending review, or expired against its deadline.
- 04Add all costs attributable to the task identifier, including discarded attempts.
- 05Divide accumulated cost by the number of accepted tasks and compare it with the deadline actually met.
A parameterized model for standard calls, caching, and batch processing
Use dollar amounts per token rather than per million within a spreadsheet: standard input e = 10/1,000,000; output s = 50/1,000,000; five-minute write w5 = 12.50/1,000,000; one-hour write w60 = 20/1,000,000; and read r = 0.25/1,000,000. For each group of requests with the same reusable prefix, define P as prefix tokens, D as average dynamic input per call, O as average output, and N as the number of calls made before the cache entry expires.
Without caching, the group cost is N multiplied by P plus D, all multiplied by e, plus N multiplied by O multiplied by s. With five-minute caching, the approximate cost is P multiplied by w5, plus N multiplied by P multiplied by r, plus N multiplied by D multiplied by e, plus N multiplied by O multiplied by s. For one-hour caching, replace w5 with w60. These formulas assume a read hit for every subsequent call and a single initial write; they are an idealized scenario, not a guarantee.
To incorporate an observed hit rate H, replace N with H multiplied by N in the read term and add the applicable input cost for calls that do not hit. The exact implementation must follow how the application groups prefixes and how the API reports usage. Different prefixes should also be separated: averaging a highly reused corpus with unique requests can hide the fact that caching works only for one portion of the workload.
For Batch, apply the documented 50% discount to the applicable components of the direct-channel model, then add the cost of resubmissions and validation. Do not assume that all jobs in a nightly queue are equivalent: jobs that expire, require priority treatment, or need corrections may end up on a different route with a different rate.
Indicative decision by observed pattern
| Condition | Mode to evaluate first | Reason and caution |
|---|---|---|
| One request or low overlap | No cache | Avoid paying for a write that may expire without a read. |
| Several close calls with the same prefix | 5-minute cache | The write is cheaper; confirm reads occur within the TTL. |
| Reuse spread across a broader window | 1-hour cache | It pays off only if it prevents enough rewrites or enables more useful reads. |
| Non-urgent, homogeneous workload | Batch, with or without cache | The discount may be meaningful, but asynchronous processing must be acceptable. |
| Long output or high rework | Measure before choosing | The context discount may be outweighed by output and retries. |
Three reproducible flows for measuring actual overlap
Case one: an engineering agent. Split the prompt into stable instructions and policies, repository state that changes at a known cadence, and the user’s one-off request. Do not assume that the entire repository should or can be included in the prefix. Measure how many tokens in the stable block repeat identically, how many actions occur within the TTL, and how many turns break the match because state changes. Also record retries caused by tool validation, failed tests, or responses that exceed the output budget.
Case two: an assistant answering over a stable corpus. A shared corpus and fixed instructions are natural cache candidates, but the analysis must distinguish between the entire corpus sent, the document excerpt retrieved for each question, and conversation history. If every question uses a different document selection, seemingly stable volume may have little exact overlap. Build groups by corpus version and prefix rather than using one global hit rate that mixes incompatible behavior.
Case three: nightly analysis. Group work that does not need an interactive response and measure the percentage that meets its delivery time through Batch. Compare invoiced savings with the cost of exceptions: urgent items removed from the batch, resubmissions, formatting errors, and jobs that expire. If the flow requires a result before later processes can start, queue time is part of its operating cost even if it is not represented as a token.
Across all three cases, the usage token counters are the source for reconstructing charges by mode. The limits documentation explains that several token categories count against the input-token-per-minute budget and provides a formula to reconstruct total input from usage data. That information is useful both for forecasting capacity and for detecting why a strategy that is cheap on paper creates waits or retries.
Minimum telemetry fields per request
| Group | Fields to record | Measurement use |
|---|---|---|
| Identity | Task ID, prefix-group ID, prompt version, corpus version | Connects attempts and identifies changes that reduce reuse. |
| Usage | Standard input, cache write, cache read, output | Calculates attributable billing and hit rate. |
| Time | Start, end, queue time, committed deadline | Separates savings from operational compliance. |
| Outcome | Accepted, technical error, retry, review, expired | Produces cost per correct task. |
| Capacity | Limit reached, concurrency, retry cause | Identifies whether limits prevent use of the selected mode. |
Limits, deadlines, and conditions that change the choice
Available capacity can invalidate a decision based on price alone. The API documentation distinguishes limits for requests per minute, input tokens per minute, and output tokens per minute. Tokens associated with caching matter to those calculations. A design that concentrates many writes or large inputs may hit limits, increase its own queue, or induce retries; in that situation, observed cost per task can rise even if the theoretical rate is low.
Batch also has documented queue limits and an expiration period. Before moving a queue, measure its age distribution and determine what share can wait without affecting dependencies. Establish an exception path for urgent jobs and budget it separately. A single queue that mixes work with different urgency makes both savings and missed deadlines difficult to attribute.
Cache hits are best effort and depend on traffic patterns, according to the Batch documentation. Treat them as a measurable outcome, not as contracted capacity. Design the application so that a missing hit remains functionally correct and so that the budget covers the scenario with the lowest reasonable reuse.
Data residency may introduce a price multiplier under documented commercial rules. If it is relevant to your environment, incorporate it into every term in the spreadsheet before comparing alternatives. It should not be selectively applied to one mode to make that mode appear more attractive.
Decision template before changing modes
- 01Select a representative week or volume and preserve segmentation by task type.
- 02Calculate cost per accepted task without caching, with a five-minute TTL, with a one-hour TTL, and in Batch where the deadline allows it.
- 03Require each alternative to clear a defined net-savings threshold after retries and validation.
- 04Also require a deadline-compliance threshold and a minimum accepted-task rate.
- 05Review unread writes, prefix changes, output distribution, and retry causes every week.
- 06Return to the previous mode or split the workload when savings disappear for a specific segment.
Conclusion: verifiable savings require segmentation and outcome control
Claude Fable 5.1 can substantially reduce the input portion of workloads that reuse a large, stable prefix and read it several times within its TTL. Five-minute caching is usually the first comparison point when requests are concentrated; one-hour caching requires evidence that its more expensive write prevents enough rewrites or enables reuse that would otherwise be lost. Batch deserves a separate evaluation for deferrable work, not an automatic conversion of the entire workload.
A robust decision does not start from one price per million tokens. It starts from real request groups, API usage fields, hit rate, time to reuse, output, retries, accepted jobs, and deadlines. With those data, an organization can set an explicit threshold: adopt a mode only if it reduces cost per correct task while keeping the service within its operational objectives.
Finally, a lower bill does not support the conclusion that the model answers with higher quality, that users perceive lower latency, or that human review decreases. Those variables must be measured with separate evaluations and metrics. Nor does it support inferring equivalent prices on Bedrock or other cloud channels. The uncertainty should remain visible in the dashboard and in any financial decision based on this guide.
Open questions
- Prices, limits, Batch conditions, and commercial multipliers may change; this guide uses the values described in the supplied sources and requires verification before contracting or deploying.
- The documentation states that cache hits are best effort; a future hit rate cannot be guaranteed from a limited test.
- No regional pricing for Amazon Bedrock or other cloud channels was supplied to make a numerical comparison between providers.
- The economic cost of human review, a missed deadline, or a low-quality result depends on each organization and cannot be inferred from token rates.
- The proposed savings and deadline thresholds are a decision method, not universal recommendations: they must be calibrated with data from the real workload.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction