Context Caching in AI: What It Reuses—and Why It Doesn’t Mean “Remembering”
01

One-sentence definition

Reutilización de prefijos de entrada ya procesados para reducir latencia o coste cuando el proveedor y la petición lo permiten.

02

Definition: reusing computations, not forming memories

In artificial intelligence, context caching is the temporary storage of processing results that a model can reuse to avoid recalculating certain parts of an input. In language models, the term usually refers to two related but distinct mechanisms: a KV cache, which preserves internal attention states during generation, and a prefix or prompt cache, which can reuse the processing of an identical or matching beginning shared across requests.

The word “context” can be misleading here. It does not mean that the system has added information to its permanent knowledge, or that it will necessarily remember a conversation after it ends. It refers to the data the model processes to produce a response, along with intermediate results an implementation may keep available for a while. A cache can reduce repeated work when the right conditions are met; on its own, it does not add new information to the response.

The specific mechanism depends on the model, runtime, and service. An implementation may expose cache-use metrics, offer controls for sharing or isolating certain data, or leave the mechanism invisible to users. “Context cache” is therefore a useful label, but it is important to clarify which type of cache is meant.

03

How a KV cache works during generation

An autoregressive model generates text one token at a time: at each step, it predicts the next token based on the preceding ones. To calculate attention, the model produces internal representations known as keys and values. A KV cache preserves these tensors for tokens that have already been processed, so later steps can reuse them instead of recalculating everything from the beginning.

In simplified terms, the model processes the prompt, calculates the required states, and stores the corresponding keys and values. Then, when it generates a new token, it calculates the states for that new step and also consults the stored states for earlier tokens. The cache grows as generation continues, within the limits of the system’s memory and management strategy.

This describes the general mechanism, not an identical data structure used by every model. There are different cache variants and strategies for managing the cache, keeping it available, or placing it at different levels of memory. A serving system may also have mechanisms for reusing blocks across requests; that should not automatically be confused with the KV cache for a single sequence.

A simplified KV-cache cycle

  1. 01The model processes the prompt and calculates internal attention states.
  2. 02The implementation keeps the keys and values for the processed tokens.
  3. 03To produce the next token, the model reuses those states and calculates the new step.
  4. 04The process repeats while generation continues and resources are available.
04

Prefix caching: reuse across requests

Prefix or prompt caching takes advantage of repeated beginnings shared by different requests. Instead of recalculating the processing for a prefix that is already available in the cache, the system can reuse the KV-state blocks associated with that prefix and continue computing from there. For example, a server may use this technique when many requests share the same initial instructions or a common section of context.

Reuse requires a sufficient match according to the implementation’s rules. In systems that organize the cache into blocks, the prefix is divided into blocks and the system checks whether earlier blocks match; it is not enough for two requests to be about the same topic. Changes to the text, its order, or its tokenization can prevent a match. Other parts of the configuration may matter too, depending on the system.

Prefix caching does not mean that a response is saved and returned unchanged. What gets reused is the processing of a shared part of the input; a new request may still contain different material and require additional computation. Whether a benefit is realized depends on the match, the availability of the blocks, and the rules for expiration or eviction.

Some providers describe this feature using specific metrics, such as the number of input tokens processed from cache. Those metrics and retention conditions belong to the documented implementation; they should not be generalized to other providers or runtimes.

What gets reused, and when

MechanismWhat it preserves or reusesTypical useMain limitation
KV cache for a generationAttention keys and values for tokens already processedContinuing a token-by-token generationThe sequence and available resources constrain its size and use
Prefix or prompt cacheKV states for a matching prefix shared across requestsAvoiding recalculation of a repeated part of the inputsIt works only if matching rules are met and the blocks remain available
An application’s persistent memoryData the system chooses to retain for future interactionsRetrieving preferences or information in another sessionThis is a separate feature with its own storage rules
05

Three practical examples

The examples below describe possible uses of the mechanism. They do not promise that a particular platform implements caching in the same way or will always produce a measurable improvement.

blocks

examples

06

Concepts that are often confused

A cache and a context window are not synonyms. The context window is the amount of information a model can consider in a given run, subject to limits set by the model or service. A cache is a mechanism for keeping or reusing computations. A cache does not, on its own, expand the context window or let you include more text than the system supports.

A cache and persistent memory are not the same thing either. An agent’s or application’s memory may store information for retrieval in later interactions, according to its design rules. A processing cache helps reuse intermediate results and may disappear, expire, or be evicted. Although both features can involve temporarily stored data, they serve different purposes.

KV caching and prefix caching do not describe exactly the same operation. The former usually refers to states maintained during the generation of a sequence; the latter refers to the opportunity to reuse states from a shared prefix across requests. A prefix-cache system may be based on KV blocks, but reuse across requests adds conditions for matching, management, and isolation.

Finally, reusing processing is not the same as reusing text as a ready-made response. If two requests share a prefix, the shared part may be able to use previously calculated internal states. That does not mean the output text has been saved, that the new response will be identical to an earlier one, or that the model has gained additional understanding.

07

Limits, privacy, and operations

A cache is not unlimited. KV states take up memory, and serving systems must manage which blocks they keep available. When resources are scarce, an implementation may evict blocks or use other management strategies; an entry that could previously be reused may no longer be available when another request arrives. Effective duration and expiration rules are not universal.

Matching matters too. A prefix that is similar in meaning may not match for caching if its token sequence differs. Changes to the model, template, or configuration can alter processing or the way prefixes are detected. So, repeating an instruction does not guarantee reuse.

For privacy, check the documentation and configuration of the specific system. Find out how long cached data is kept, how requests or users are isolated, whether partitioning or invalidation controls are available, and what metrics are exposed. Sharing a cache across requests can raise isolation considerations; for example, the documentation for some systems warns of side-channel risks and describes specific mitigations. That warning does not show that all services have the same risk, or that a particular mitigation is available everywhere.

Metrics also need context. A counter for tokens served from cache may describe input reuse, but on its own it cannot show how much cost or latency was saved. Those outcomes depend on the API, hardware, workload, configuration, and billing or measurement method. Quantitative claims should be attributed to the provider and the conditions under which they were measured.

A checklist for evaluating a cache

  1. 01Identify whether you are dealing with a per-generation KV cache, a prefix cache across requests, or both.
  2. 02Check what must match for states to be reused and which changes invalidate a match.
  3. 03Review documented memory limits, expiration, eviction, and isolation controls.
  4. 04Distinguish cached-token metrics from measurements of latency, cost, or throughput.
  5. 05Check cache-retention policies separately from conversational-memory or application-storage policies.
08

Related concepts and practical guidance

Attention is the model mechanism that relates different parts of a sequence; keys and values are internal components used in that calculation. Tokens are the input and output units the model works with. Inference is the process of using the model to produce an output. A KV cache helps avoid repeating part of the attention work, while a prefix cache aims to reuse the processing of a shared input. These concepts are connected, but they are not interchangeable.

When reading documentation or configuring a system, first ask what is being stored: KV states, a processed prefix, conversational text, or something else. Then find out where the cache applies—a single generation, a session, or multiple requests—how long it lasts, what makes a match valid, and how the system handles memory pressure. Finally, separate what the platform measures from what you expect: reusing computations may reduce repeated work, but the actual savings and latency depend on the circumstances.

As a general guide, if you want to continue a token-by-token generation, look for information about KV caching and its strategies. If several requests share a long opening, look for documentation about prefix caching and its matching rules. If you need the system to retrieve user data across sessions, consult the application’s persistent-memory or storage feature; do not assume a context cache will do that.

A quick decision guide

NeedConcept to look upPractical question
Avoid recalculating states during an ongoing generationKV cacheWhat caching strategy and memory limits does the runtime use?
Reuse a shared prefix across callsPrefix or prompt cacheWhat needs to match, and how long do the blocks remain available?
Retrieve preferences or data in future sessionsPersistent memory or application storageWhat is stored, who can access it, and how can it be deleted?
09

Quick examples

10

Related concepts

11

Sources consulted