Ilustración editorial para Cuantizar un modelo local sin adivinar: cómo elegir 4, 6 u 8 bits según VRAM, contexto y pérdida aceptable
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

Weights fitting in VRAM does not mean the system will work

The most common mistake when selecting a local model is to compare the size of the quantized file with GPU memory and consider the decision settled. That calculation only covers the weights, approximately. During inference, the system also uses the key-value cache—the KV cache—activations and temporary buffers, runtime-reserved space, the context for each request, and, in a service, simultaneous requests. A configuration can load the model, answer a short question, and still fail when it receives a long document or several requests at once.

The practical consequence is important: “fits” should mean that the configuration completes the expected maximum load with measured headroom, not that it starts an isolated session. If that headroom disappears, the result may be an out-of-memory error, an automatic reduction in context, moving part of the workload to system RAM, or highly uneven latency. Which behavior occurs depends on the runtime and its configuration; it should not be assumed without verification.

Quantization is one of several levers. Reducing weight bit width usually frees memory and may allow a larger model, but it does not by itself eliminate the growing cost of the KV cache as context length or concurrency increases. It can also change quality, performance, or the available execution paths depending on the format and backend. That is why there is no universal equivalence between “4-bit,” “6-bit,” and “8-bit.”

This guide starts from a concrete workload: the input length that must be accepted, the number of tokens to generate, the number of requests that will coexist, the latency that remains useful, and the errors that would be unacceptable. If you are still selecting a model family, first review the local models guide. If the question is between different base models, use the comparison section before attributing differences caused by the model itself to quantization.

02

The memory components that must be separated

A useful estimate begins by breaking memory down into components that can be observed separately. The first is model weights. Their size depends on the parameter count, the quantized representation, and format-specific metadata, such as scales, blocks, or auxiliary structures. Simply dividing parameters by eight, six, or four therefore provides orientation, but it does not replace the actual size reported by the format and runtime.

The second component is the KV cache. In an autoregressive decoder, the system retains keys and values for already processed tokens so that it does not recalculate them for every generated token. Transformers documentation describes cache tensors with batch, heads, sequence-length, and head-dimension axes. There are both keys and values, and this storage is repeated for every layer. For the same architecture, increasing context, batch size, or concurrency increases the memory required.

The third component includes activations and temporary buffers. Their size depends on the backend, compute precision, kernels, prefill for long inputs, generation length, and how requests are grouped. They should not be replaced with a universal constant. The fourth component is memory not directly assigned to the model, such as the execution context, libraries, allocators, and fragmentation. The fifth is an explicit operating margin: memory deliberately left unallocated to tolerate spikes, measurement differences, and real load.

On a server, batch size and concurrency require an additional distinction. A batch may be the number of sequences processed together in one step, while concurrency is the number of live requests. Depending on the scheduler, these values may be related but are not interchangeable. To estimate cache use, count the sum of live tokens across sequences that coexist, not only the maximum length of one request.

Components to record before deciding

ComponentWhat determines itHow to verify it
WeightsBase model, format, and quantizationMemory after model loading or runtime report
KV cacheLayers, KV heads, head dimension, live tokens, dtypeCache capacity or usage and effective served length
Activations and temporariesPrefill, generation, batch, kernels, and backendPeak memory during representative load
Runtime memoryLibraries, allocator, device context, and fragmentationMemory before and after starting the process
HeadroomVariability and expected maximum loadMinimum free memory observed in repeated tests
03

Estimate weights and KV cache before downloading or deploying

The estimate is not meant to predict every byte. Its purpose is to rule out infeasible configurations and determine which tests are worth running. For weights, use the size reported for the specific artifact you intend to load, not a generic figure for the model family. If only the parameter count is known, treat the result as an incomplete theoretical minimum. Quantization formats store additional information, and some runtimes convert or duplicate structures during loading.

For a Llama-like architecture, a conceptual approximation for KV cache per sequence is: layers multiplied by two, multiplied by the number of key-value heads, multiplied by sequence length, multiplied by head dimension, and multiplied by bytes per cache element. The factor of two represents K and V. For multiple simultaneous sequences, add the resident tokens from each one. If the runtime uses a static cache, it may reserve capacity up to a maximum even when instantaneous use is lower; if it uses a dynamic cache, usage may grow with the request. Both strategies require measurement.

It is essential to use the number of key-value heads, which is not always the total number of attention heads. In multi-query or grouped-query attention, several query heads share KV projections. This can substantially reduce cache use compared with an architecture with one KV projection per attention head. The exact model configuration must provide the layer count, KV-head count, and head dimension; these should not be inferred from a commercial model name.

The bytes per cache element must also be verified. Quantizing weights does not automatically mean quantizing the KV cache. Some environments allow a quantized cache data type to be selected, others use a different precision by default, and others apply offloading. These options change the memory budget and may affect performance or numerical behavior. Record the runtime’s actual configuration, not only the bit width in the file name.

04

What really changes when moving from 8 to 6 or 4 bits

As a general rule, reducing weight precision reduces the memory footprint relative to a higher-precision representation of the same model. That can make a smaller GPU viable, leave more budget for context, or support more concurrent requests. However, the observed reduction does not necessarily follow an exact eight-to-six-to-four proportion. Packing, per-group scales, file format, internal conversions, and backend buffers all change the outcome.

Quality does not depend on bit width alone either. The quantization algorithm, group size, tensors receiving special treatment, base model, and task all matter. A 4-bit quantization from one method may preserve a task well while a 6-bit quantization from another method may not; the reverse can also happen. For that reason, bit width is useful for forming test hypotheses, not for certifying accuracy.

Speed requires the same caution. Lower memory use can reduce transfers and improve viability on a constrained device, but a format may lack efficient kernels on a specific backend or require conversions. Increasing context can shift the bottleneck toward cache management and prefill. Performance should be measured with two separate metrics: time to first token for representative inputs and subsequent generation rate. A single tokens-per-second figure hides important differences.

As a starting point, test 8 bits when quality is critical and the budget permits it; test 6 bits when a meaningful amount of memory must be recovered without immediately moving to the most aggressive option; and test 4 bits when VRAM is the dominant constraint or tests show no unacceptable loss. These are testing priorities, not universal recommendations.

Condensed decision tree

Observed situationFirst actionWhat not to assume
Weights do not fit with headroomTry lower precision or a smaller modelThat lowering bits will solve context cost
Weights fit, but long inputs failReduce target context, inspect KV cache, or use more VRAMThat file size predicts context capacity
It fails with several requestsSize for concurrent live tokens and actual batch behaviorThat a single-session test represents the service
Quality drops on critical tasksIncrease precision, change method, or use a smaller model at higher precisionThat more parameters compensate for any loss
Latency is unstableMeasure prefill, generation, offloading, and free memoryThat average tokens per second is enough
05

Decision procedure: from constraint to viable candidate

Define the operating contract first. Write down the maximum input length that must genuinely be supported, an output-token reserve, the maximum number of live requests, the latency target, and critical tasks. Distinguish an exceptional maximum from the usual target. If an application processes long documents, measuring only short messages does not represent either its memory risk or its useful quality.

Next, collect architecture and runtime parameters. For the model, record layers, KV heads, head dimension, and the weight format. For the runtime, record cache dtype, whether the cache is static or dynamic, whether offloading is available, the GPU memory limit, and any batch or in-flight-token settings. In serving tools, the cache budget may be set explicitly or derived from a fraction of available memory; both cases must be documented in the experiment.

Calculate a range, not a single figure: observed or estimated weights, KV cache for the target load, a reserve for temporaries, and headroom. If the total exceeds available VRAM before headroom is applied, reject the combination. If it fits only narrowly, classify it as a high-risk candidate and test it under maximum load. If it fits with headroom, do not approve it until quality and latency have been validated.

Choose at least three candidates that test distinct hypotheses: the desired model at 8, 6, and 4 bits; or, when one candidate makes no sense, a smaller model at higher precision. Keep the base model, revision, prompt, maximum context, output limit, seed where supported, decoding parameters, hardware, and runtime version constant. Changing multiple variables at once makes it impossible to attribute a difference to quantization.

A reproducible seven-step process

  1. 01Define input context, reserved output, concurrency, and target latency.
  2. 02Record the architecture, weight artifact, and cache configuration.
  3. 03Estimate weights, KV cache, and headroom for the maximum number of live tokens.
  4. 04Reject candidates that do not fit before headroom or require unverified assumptions.
  5. 05Run comparable candidates with identical parameters.
  6. 06Measure memory, time to first token, generation, errors, and output quality.
  7. 07Keep a configuration only if it meets the quality threshold and retains headroom under maximum load.
06

The minimum test for detecting a loss that matters

A useful test does not need to be huge, but it does need to be representative. Build a small set of cases that includes the work motivating deployment: structured extraction, classification, document synthesis, coding assistance, or constrained answers, as appropriate. Include inputs of typical length and some close to the operating limit. Long inputs are necessary because they can reveal both memory failures and losses in instruction following or detail retrieval.

For every case, define what will be validated before running the model. Some tasks support exact comparisons: valid JSON matching a schema, permitted labels, required fields, a query that must contain specific values, or automated tests for code. Others require human review using a rubric: fidelity to the document, coverage, lack of invented content, format compliance, and usefulness. Do not confuse fluency with correctness.

Set an explicit threshold. For example, a configuration may be rejected if it fails more critical cases than the reference candidate, if it worsens structural validity beyond a team-defined limit, or if it introduces new errors in sensitive data. The threshold belongs to the application’s risk profile; it cannot be inferred from bit width. A creative drafting task can accept more variation than data extraction for a downstream process.

Repeat the tests. With stochastic decoding, multiple runs help distinguish generation variation from systematic degradation. With deterministic decoding, repeats are still useful for observing performance stability and memory errors. Report results by task type and input length, not only as a global average. A quantization that appears equivalent on average may concentrate its failures in the longest documents or the highest-impact task.

07

Three operating profiles and their priorities

On a laptop with a limited GPU, the priority is usually avoiding a configuration that continuously depends on system RAM or offloading for an interactive experience. Start with a realistic context, one request, and a model or quantization that leaves headroom. If 4 bits is the only way to load the model, also compare a smaller model at 6 or 8 bits. The latter may be more stable and more useful for the specific task despite having fewer parameters.

On a workstation with one GPU, there is more room to choose among quality, context, and speed, but the limit is still shared by weights, cache, and temporaries. This is an appropriate environment for comparing 4, 6, and 8 bits on the same corpus and deciding whether VRAM should be reserved for long contexts. If you expect to switch between brief sessions and document analysis, measure both profiles: the result of a short conversation does not size the latter.

On a server with moderate concurrency, the planning unit is no longer the model file but total live-token capacity. The cache manager and request scheduler become part of the decision. A configuration that works for one session may exhaust memory when long prefills overlap. Define admission limits, maximum length, output reserve, and concurrency; then test bursts and mixtures of short and long requests. Queue metrics and latency percentiles are more informative than the best isolated result.

In all three profiles, offloading is an option that must be declared, not an invisible solution. It may expand apparent capacity by moving part of the data, but it can also alter latency and depend on the connection between CPU and GPU. Use measurements to decide whether that tradeoff is acceptable for the use case.

08

Signals to reject a configuration and an adoption checklist

Reject a configuration when it produces intermittent memory errors, even if a short demonstration works. Intermittence usually indicates that available memory depends on request shape, prefill peaks, fragmentation, or other process loads. It is also a reason to reject the configuration when the runtime silently reduces effective context, when the system is stable only under a lower load than planned, or when observed headroom disappears in repeated tests.

Quality degradation must be analyzed by pattern. Errors concentrated in extraction, calculations, mandatory fields, instruction following, or long documents carry more weight than stylistic changes if those tasks are critical. An apparently reasonable answer with invented values should not be approved solely because of an average score. Also review format validity whenever the output feeds downstream software.

Before setting a default option, check privacy and operations. Running inference on the machine does not itself guarantee that no data leaves it: model downloads, telemetry, prompt logging, dependency updates, and observability tools are separate concerns. What data each component retains and what communications it performs must be verified in the selected configuration and network environment.

The final result does not have to be “the highest quantization possible.” It may be 6 bits for an interactive profile, 4 bits for a large-context analysis profile, or 8 bits for a task where the detected loss is unacceptable. Keep the alternatives with their test records. If the runtime, hardware, weight format, or workload changes, measure again: the previous conclusion is no longer a guarantee.

Checklist before adopting a quantization

  1. 01Weights, KV cache, temporaries, and headroom have been measured or justified separately.
  2. 02Context, reserved output, and concurrency reflect the expected maximum load.
  3. 03It has been verified whether quantization affects weights, KV cache, or both.
  4. 04Compared configurations use the same base model and equivalent conditions.
  5. 05The corpus includes critical tasks and representative long inputs.
  6. 06An acceptable-loss threshold was defined before reviewing results.
  7. 07Minimum free memory, errors, and latency percentiles have been recorded.
  8. 08The offloading, telemetry, download, and logging configuration has been reviewed.
  9. 09The decision can be reproduced from the complete technical record.

Open questions

  • The KV-cache formula presented here is a conceptual approximation. Actual allocation may vary because of static or dynamic caching, paging, alignment, buffers, and the runtime’s attention strategy.
  • An acceptable quality loss cannot be determined without knowing the task, critical errors, and the team’s validation procedure.
  • The effect of 4, 6, or 8 bits on latency and quality depends on the quantization format, model, available kernels, and hardware; it requires local measurement.
  • The availability and exact meaning of cache, offloading, and memory-budget options change across runtime versions.
09

Keep exploring

09

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction