
One-sentence definition
Reducción de la precisión numérica de pesos o activaciones para usar menos memoria y acelerar la inferencia.
Definition: Representing Numbers with Fewer Levels
Quantization in AI is the process of representing numerical values with a smaller set of levels than the original representation uses. In a neural network, those values may be weights, intermediate activations, or elements of a language model’s key-value (KV) cache. Using fewer levels can reduce the memory required and, when the software and hardware support suitable operations, can also reduce computation cost or time.
This definition describes a numerical transformation; it does not, by itself, guarantee that a model will take up an exact fraction of its original size, run faster, or retain its quality unchanged. The outcome depends on which components are quantized, how they are encoded, and how the implementation uses them. So saying that a model is “4-bit” provides incomplete information, not a performance prediction.
This article separates two questions: which values are transformed, and how those values are mapped to the available levels. The first determines the scope of quantization; the second determines the representation and its potential error. For related vocabulary, see the glossary entries on quantization, inference, local models, and LoRA, as well as the comparison section.
How the Mapping Works
In linear quantization, a scale relates integer levels to values in the original domain. Many schemes also use a zero point, which indicates which integer code represents real zero; this can matter, for example, when values are not distributed symmetrically around zero. The process usually involves rounding and, if a value falls outside the representable range, clipping it to the nearest available endpoint. When reconstructing the value, the scale and zero point are used to return an approximation in the original domain.
There is no single convention used by every method. Some formats use scales and zero points; others use codes or non-uniform levels. Scale, zero point, and quantization code are therefore related but not interchangeable concepts: each describes a different part of the representation. The way parameters and ranges are selected, and how values are grouped, can also vary.
Elements that Describe a Quantization Scheme
These terms identify different decisions within a quantization scheme.
| Element | What it describes | Why it matters |
|---|---|---|
| Number of bits | How many bits encode each value in a particular part of the model. | Limits the available codes, but does not summarize all storage costs. |
| Scale | The relationship between codes and real values. | Affects resolution and reconstruction error. |
| Zero point | The code that represents real zero in some linear schemes. | Can support ranges that are not symmetric around zero. |
| Granularity | Whether parameters are shared per tensor, channel, or block. | Changes how well the representation adapts and how much metadata is needed. |
| Calibration | The procedure and data used to estimate parameters or ranges. | The outcome may depend on how representative those data are. |
What the Bit Width Does—and Does Not—Mean
With b bits, up to 2 raised to the power of b distinct patterns can be encoded. A 4-bit scheme therefore has up to 16 possible codes per value. But the number of codes does not, on its own, determine which values they represent, the effective precision achieved, or the storage size of the complete model. The way levels are distributed and the use of metadata such as scales and zero points also matter.
Actual storage may include auxiliary information and packing structures in addition to the quantized values. A model in operation also needs memory for components that may not be quantized in the same way, such as activations, KV cache, program buffers, or other runtime components. The proportion of the model that remains at higher precision also changes the total.
Speed cannot be inferred from bit width alone either. To achieve acceleration, the hardware, kernels, and runtime must efficiently support the representation and the relevant operations. In some environments, conversion or data-access overhead can add costs; in others, moving less data can be beneficial. A useful comparison is empirical and should be made with the implementation of interest.
Which Parts of a Model Can Be Quantized
Weight quantization transforms the network’s learned parameters. It is a common option for reducing the storage needed to load models for inference. It does not mean that the whole model has been quantized: activations produced during computation may remain at a different precision.
Activation quantization represents the intermediate values flowing through a network with fewer levels. It can be combined with quantized weights, but it does not have to be. For example, SmoothQuant studies post-training quantization of both weights and activations in large language models (LLMs). Its results apply to the method and published evaluations, not as a guarantee for every model or device.
In generative language models, the KV cache stores intermediate information associated with tokens that have already been processed. Quantizing this cache is different from quantizing weights: it affects another component of memory and can have different effects on inference. The KVQuant paper focuses specifically on KV-cache quantization, illustrating why it is useful to state precisely what is being transformed.
A reproducible description of a configuration should avoid saying only “quantized model” without specifying what was quantized. At a minimum, distinguish weights, activations, and KV cache, and say whether unmentioned components retain higher precision.
Components and Practical Questions
This table helps clarify the scope of a configuration; it does not imply that every method quantizes all three components.
| Component | What is represented with fewer levels | Question to ask during evaluation |
|---|---|---|
| Weights | Learned parameters. | How much memory do the weights use, and what precision do they retain? |
| Activations | Intermediate values produced during computation. | Which operations and layers are involved, and what impact does this have on the task? |
| KV cache | Stored state used to generate sequences. | How does it change memory use and behavior in the context being evaluated? |
When Quantization Happens: PTQ and QAT
Post-training quantization, or PTQ, is applied after a model has been trained. Depending on the method, it may use calibration data to estimate ranges or other quantization parameters. These data do not necessarily serve to retrain all the weights; they help configure the mapping. As a result, the choice and representativeness of calibration data can affect the outcome.
Quantization-aware training, or QAT, incorporates a simulation or consideration of quantization effects during training so that the model can adapt to them. The work by Jacob and colleagues presents a method for quantization and training aimed at inference using integer-only arithmetic. PTQ and QAT differ in timing and procedure; there is no basis for claiming that one is always better. The right outcome depends on the network, objective, and available process.
Granularity matters too. A scheme may share quantization parameters across an entire tensor, use parameters per channel, or group values into blocks. More localized granularity may describe differences in value distributions more closely, but it also comes with its own costs and requirements. Two configurations with the same nominal bit width but different granularity should not be assumed equivalent.
How to Read a Quantized Configuration
Before comparing two results, identify the components and conditions they describe. If an important detail is missing, record it as an uncertainty instead of inferring it from the bit width.
- 01State whether the method is PTQ or QAT and, where relevant, what calibration data were used.
- 02Specify whether weights, activations, KV cache, or another component were quantized, and what was left out.
- 03Record bit width, representation type, granularity, and relevant metadata.
- 04Separate weight memory from total memory during execution.
- 05Measure quality and performance on the task, hardware, and runtime where the model will be used.
Three Applied Examples
The following examples show different uses of quantization; they are not format recommendations or promises about results. In each case, the effect should be checked with the actual model, implementation, and usage conditions.
The examples illustrate why it helps to state both the component being quantized and the goal of the deployment. Reducing weight storage, limiting memory on a vision device, and fine-tuning a quantized model are distinct use cases, even though all involve quantization.
Related Concepts That Should Not Be Confused
Pruning removes or reduces connections, weights, or structures to make a network sparser or more compact. Quantization, by contrast, changes the numerical representation of the values that remain. Methods can combine the two ideas, but they are not the same procedure.
Distillation transfers behavior from a teacher model to a student model through a training process. It can produce a smaller model, but its mechanism is not mapping values to fewer levels. A distilled network can also be quantized afterward: these are compatible but distinct choices.
LoRA is a fine-tuning technique that adds low-rank adapters to modify a model’s behavior. QLoRA combines adapter fine-tuning with a quantized base. Thus, LoRA is not a quantization format, and QLoRA does not mean any low-precision model.
INT4 usually refers to a four-bit integer representation, whereas NF4 is a four-bit data type proposed in the QLoRA paper to represent weight values. They are not interchangeable names: they describe different representations. In addition, the nominal data width does not, by itself, reveal the implementation, packing, or total cost.
Finally, reducing precision is not the same as changing from one floating-point precision to another without changing the format’s purpose. Reduced-precision formats can have different ranges and properties. “Lower precision” is a broad description, whereas “quantization” refers to a mapping process that should be specified.
The Difference in One Sentence
These categories can be combined in a workflow, but they describe different operations.
| Term | Main operation |
|---|---|
| Quantization | Representing values with fewer levels. |
| Pruning | Removing or reducing parts of the network. |
| Distillation | Training a student model using signals from a teacher. |
| LoRA | Fine-tuning a model using low-rank adapters. |
| QAT | Training while taking quantization effects into account. |
Limits, Common Mistakes, and Practical Criteria
Quantization error arises because an original value may not have an identical representable level. The size and distribution of these errors can affect a network’s output, and not necessarily affect every layer or task equally. A relevant evaluation should examine the behavior that matters for the application, not just a generic metric or the format label.
A common mistake is to assume that fewer bits guarantee a proportional reduction in memory. Metadata, packing, components that remain at higher precision, and runtime memory all change the total. Another mistake is to conclude that fewer bits always speed up inference: if the runtime or kernels do not take advantage of the representation on the available hardware, latency may not improve, or may be affected by additional overhead.
It is also worth avoiding the assumption that calibration is an irrelevant detail. If a method depends on representative data to estimate parameters, calibration data that poorly match real inputs may not reflect the intended use adequately. Model sensitivity, input distribution, context length, and task can all change the outcome. Papers on GPTQ, AWQ, SmoothQuant, and KVQuant evaluate specific methods and configurations; their published figures should not be extrapolated without measurement.
When documenting a configuration, record the method, which components are quantized, precision and granularity, calibration data where relevant, observed memory, quality metrics, hardware, runtime, and test conditions. This makes it possible to compare results without turning a label such as “4-bit” into a general claim about quality or speed.
Practical Criteria Before Comparing
A useful comparison keeps relevant conditions consistent and makes explicit the details that could change how the results are interpreted.
- 01Define the problem and the inputs the model is expected to handle.
- 02Specify which components are quantized and which retain another precision.
- 03Record the method, bit width, granularity, calibration approach, and format.
- 04Measure weight memory and peak runtime memory separately.
- 05Evaluate quality, latency, and performance with the target hardware and runtime.
- 06Report limitations and missing information; do not present one test as a universal guarantee.
Where to Go Next
Quantization is a representation tool, not a category that by itself describes a system’s complete behavior. When reading a technical specification or evaluation, start by identifying what was quantized, how the values are encoded, and the environment in which results were measured.
For a conceptual path through the topic, consult the glossary index and related entries on inference, local models, LoRA, distillation, and evaluation. The comparison section can help organize differences between methods when comparable conditions are available; this article does not recommend a universal format or replace an applied evaluation.