Ilustración editorial para ¿Cabe este modelo en tu GPU? Cómo estimar la VRAM antes de elegir hardware
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

The question is not just how many parameters the model has

Knowing a model’s parameter count can help you get your bearings, but it does not answer the practical question: will it run with the configuration I need on this GPU? To answer that, you need to consider the specific model, its weight format, the runtime, the context you want to process, simultaneous requests, and the memory already in use by the system.

That is why “fits” is not an isolated property of a model. It may fit when loaded and run out of memory as the context grows; it may work for one request but not several; or it may start because some of the work is being done in RAM, even though the entire model is not on the GPU. The result can also change when you switch runtimes or options.

This guide will help you build an initial estimate and then test it under controlled conditions. An estimate can rule out configurations that are clearly infeasible and help you decide what to measure. It cannot replace a test using the actual hardware, runtime, and workload. Nor can it tell you the speed, response quality, or long-term stability you will get.

02

What uses memory during inference

A useful estimate separates at least four components: model weights, KV cache, temporary compute buffers, and the space required by the runtime together with the system and other processes. This is a conceptual breakdown. Depending on the runtime, its options, and the hardware, allocations may be reported differently, and monitoring tools may not show them as separate categories.

Weights are the model data the runtime loads to perform inference. The file size can offer a clue, but it does not automatically equal the total memory required. The format and quantization affect the size of the weights, and loading also involves working memory and runtime structures. Do not treat the file size as the amount of VRAM the model will use without measuring the chosen configuration.

The KV cache stores information associated with tokens that have already been processed so generation can continue. Its footprint depends on the active workload and model characteristics, not just the model’s name or parameter count. The intended context and number of parallel requests are variables to set before estimating; the architecture and configured cache type can also change the result.

Compute buffers are temporary memory used during operations in the runtime. Their size may depend on the implementation and enabled options. The runtime and system also need headroom, and the GPU may share memory with a display or other processes. Therefore, do not treat a card’s nominal memory capacity as if it were all available to the model.

Initial memory inventory

Record what you know and what you still need to check. These categories help organize an estimate, but do not imply that the runtime exposes each use separately.

ComponentWhat affects itWhat to record
WeightsModel, format, and quantizationExact size and format of the artifact loaded by the runtime
KV cacheActive context, concurrent requests, architecture, and cache strategyContext, concurrency, and cache configuration
Compute buffersRuntime, operation, and execution optionsObserved peaks during loading and inference
Runtime and systemAdditional processes, display memory, and system configurationFree memory before and during the test
03

Gather the inputs before estimating

Start by identifying the exact model artifact: not just its commercial name, but also the file or format you plan to load and its associated configuration. A configuration file may reveal architectural details the name does not, such as the number of layers and certain attention dimensions and head counts. These details help describe the architecture, but they are not enough by themselves to calculate final memory use: how the runtime represents and allocates the cache and buffers matters too.

Next, specify the runtime and its options. Record the version or configuration you are testing, the requested context length, the expected concurrency, and any choice about cache type or location. If the runtime lets you put some layers on the GPU, offload part to the CPU, or distribute the work across cards, record those choices as well. An estimate that omits these conditions mixes different scenarios.

Finally, record the system’s starting point: total and available GPU memory, processes already using it, and whether the card is dedicated to inference or has other jobs. GPU management tools may report total, reserved, used, and free memory. These are observations of the device’s state, not a complete explanation of which model component occupies each block.

Configuration checklist

Fill this out before estimating. Keep the same values during the initial test so you can attribute changes to a specific variable.

  1. 01Identify the model and the exact weight format the runtime will load.
  2. 02Record the architecture parameters available in the model configuration; mark any undocumented parameters as unknown.
  3. 03Specify the runtime and execution options, including any offloading to RAM or splitting across GPUs.
  4. 04Define the maximum context you actually expect to use and how many requests may be active at once.
  5. 05Measure free GPU memory before starting and record the processes already using it.
04

How to build an initial estimate

As a first approximation, think of required VRAM as the sum of weights resident on the GPU, KV cache stored on the GPU, buffers, and runtime overhead, plus headroom for workload variation and other uses of the card. This is not an exact formula or a figure you can fill in using generic data: the categories and their sizes depend on the implementation and configuration.

Separate known values from estimates. File size is observable, but it does not prove how much of that size will be resident on the GPU or what the memory peak will be. Your intended context and concurrency are requirements you choose. The architecture and cache strategy require information about the model and runtime. Buffers and headroom often need to be measured on the system where the model will run.

If an important variable is unknown, do not hide it behind a single number. It is more honest to prepare scenarios: for example, one with a moderate context and another with a more demanding context, each with the concurrency your service needs. Those labels do not guarantee specific memory use; they help ensure the test covers real operating conditions instead of validating only the easiest case.

Runtime documentation can help you understand the available options and stated limits, but a declared capacity or configuration value does not prove that the system can sustain the workload in practice. Use the documentation to design the test, and use measured results to check the configuration.

05

The KV cache changes with context and workload

The KV cache deserves particular attention because its footprint is tied to active work. A longer context may require retaining more information to continue generation; several simultaneous requests may keep several sequences active. You cannot infer a precise figure from the model’s name or parameter count.

Architecture matters. A more grounded estimate requires model configuration attributes and details about how the runtime handles attention and the cache. Even with that information, the observed value may depend on the selected cache type, memory allocation, and runtime options. A generic rule that lacks those inputs can provide qualitative guidance, but it cannot guarantee that a configuration will fit.

Runtimes may offer different strategies, such as dynamic, static, quantized, or CPU-offloaded caches. Changing strategies can change memory placement and execution conditions. Do not compare two estimates as if they were equivalent if they use different strategies, runtimes, or runtime options.

What to check as the cache grows

Use this table to investigate what may have caused an observed change; it does not assume a universal growth rate.

Test changeWhat may be changingWhat to hold constant or record
Increase contextMore active tokens and a different cache allocationRequested context and memory during loading and generation
Increase parallel requestsMore active sequences and more cache associated with the workloadNumber of simultaneous requests and test duration
Change cache strategyDifferent representation, location, or memory allocationCache type and exact runtime options
Change runtimeDifferent implementation and memory managementRepeat the test; do not simply carry over the previous result
06

Full GPU execution, RAM offload, or multiple cards

An execution may use the whole GPU for the model, place only some layers on it, offload part of the cache to the CPU, or split the model across GPUs. Some of these options make it possible to try configurations that would not fit with all their components on a single card, but they do not prove that the model is fully resident in VRAM. Nor do they, by themselves, tell you what performance to expect.

Check the execution mode in the runtime options and logs. In tools that let you specify GPU layers or split work across cards, record the values actually applied. If CPU offload or a CPU cache strategy is enabled, GPU memory alone no longer represents the execution’s total memory use. Distinguish “the application started” from “the configuration meets the defined residency and workload requirements.”

With multiple GPUs, knowing the sum of their nominal memory capacities still does not tell you how the model will be distributed. The split depends on the runtime’s capabilities and the chosen configuration. Evaluate each device and the effective allocation; do not assume all aggregated memory is available for every distribution.

What it means when the process starts

Classify the result based on the execution mode you verified, not merely on whether the process started without an error.

ResultWhat you can concludeWhat you cannot conclude
Weights and cache on the GPU as specifiedThe observed execution uses the GPU according to the verified optionsThat it will support every context length, concurrency level, or duration
Some layers or cache on the CPUThe runtime continued by offloading or distributing memoryThat the entire model fits in VRAM
Model loaded, but intended workload not testedLoading completed under those conditionsThat long context or several requests will work
Failure during loading or testingThe current configuration did not complete that runThat the model cannot work with other options or hardware
07

Validate the estimate with a controlled test

The test should reproduce the scenario you intend to deploy. Fix the model, format, runtime, context, concurrency, and cache and offload options. Record the initial memory state and any unrelated process sharing the GPU. If you change several options at once, it will be difficult to tell which one explains the result.

Load the model and record used and available memory. Then run a request with the intended configuration. Increase context or concurrency gradually, changing one variable at a time, and note when an error appears, offloading is activated, or usage approaches the observed limit. Do not treat one reading as a stable maximum: monitor memory during execution, because use can differ between initial loading and inference.

Use runtime logs to verify how work is distributed between GPU and CPU and which options were actually applied. Compare those observations with a GPU monitoring tool that reports total, used, reserved, and free memory. That reading helps characterize the device state; by itself, it may not separate weights, cache, and buffers. Repeat the run to detect variation on the same machine, and record the exact conditions.

Define in advance what it means to pass: complete the intended context, support the required concurrency, and avoid relying on offload that was not part of the plan. If the system needs headroom for other processes, include that as a requirement. A test that finishes without errors but leaves no such headroom may still be inadequate for the intended use.

Verification procedure

Follow this sequence without changing several conditions at once. Keep the logs so you can repeat the test after changing hardware or runtime.

  1. 01Record the model, format, runtime, cache options, context, concurrency, and distribution mode.
  2. 02Measure free memory before loading and record other processes using the GPU.
  3. 03Load the model and inspect memory and runtime logs; confirm whether any layers or cache are on the CPU.
  4. 04Test a small workload first, then increase context while keeping concurrency fixed.
  5. 05Reset or restore a comparable state, then increase concurrency while keeping context fixed.
  6. 06Record errors, observed peaks, RAM offload, and the outcome of each run.
  7. 07Repeat the test and assess whether the available headroom meets the defined operational requirement.
08

Common mistakes and limits of the estimate

The most common mistake is to use file size as if it were total memory use. That value does not necessarily include the cache, buffers, runtime, or memory already used by the system. Another mistake is to compare against the GPU’s nominal capacity without measuring how much is available in the machine’s actual operating state.

It is also easy to test only whether the process starts and assume that this validates a long context or several requests. Initial loading and sustained workload are different stages of the test. If you change context, concurrency, cache type, runtime, or the split between GPU and CPU, you are testing a different configuration.

Finally, be cautious about extrapolating a figure from one runtime to another. Tools can differ in how they manage memory and expose metrics. Observed figures describe the hardware and options tested; they do not guarantee the same result on another system or indefinitely stable execution. The test also does not establish output quality or sufficient speed for a particular use case.

Practical decision guide

An estimate helps determine what to do next. If the evidence does not answer a critical question, the right conclusion is “needs testing,” not “fits.”

SituationReasonable decisionNext step
The estimate clearly exceeds available memory before adding cache and buffersRule out that configuration on this GPU or change the requirementsEvaluate another format, distribution, or hardware, then measure again
Initial loading succeeds, but the required context has not been testedDo not consider the use case validatedIncrease context in a controlled way
The process works through unplanned CPU offloadDo not claim the model fits entirely in VRAMCheck runtime options and decide whether offloading is acceptable
Required context and concurrency pass repeated tests with headroomThere is practical evidence for this configuration on this hardware and runtimeDocument the conditions and repeat after changing components
09

Final criteria for choosing or reusing hardware

Before buying a GPU, identify a representative workload and check whether the available memory can hold the components you want to keep on the GPU, along with the headroom the system needs. If your estimate already clearly exceeds usable capacity, there is no need to pretend to have precision: that combination requires changing the requirements, distribution, or hardware. If it is close to the limit, testing on the exact system is especially important.

When reusing a GPU, measure its actual state and check what else shares its memory. Do not treat total capacity as free capacity. If you accept partial execution in RAM or distribution across cards, record that choice as part of the configuration, not as an invisible detail. When comparing options, use the same workload and criteria.

To explore local models further, see the local-model guide, the comparison tool, and the discovery section. Apply the same discipline when evaluating an option: identify the exact configuration and check the conditions that matter for your use case. The useful conclusion is not a universal VRAM figure, but a reproducible result for a defined model, runtime, hardware setup, and workload.

Open questions

  • The cited documentation does not provide a universal formula for calculating total VRAM use for every model, runtime, and hardware setup.
  • How weights, KV cache, buffers, and runtime memory are separated in observations depends on how the runtime manages and exposes its allocations.
  • Usage can vary between runs and runtimes; tests should be repeated on the intended machine with the intended options.
  • Model configuration parameters alone are not enough to infer final memory use without knowing the runtime’s strategy.
10

Keep exploring

10

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction