Ilustración editorial para Servidor local de IA: cómo calcular la concurrencia real antes de prometer un asistente privado para todo el equipo
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

The demo figure is not service capacity

A model that generates text quickly for one person does not, for that reason alone, have proven capacity for a team. An isolated demonstration often starts with a short request, no queue, a warm cache and process, and no other conversations holding memory. A shared service runs under different conditions: requests arrive at irregular rates, have heterogeneous inputs and outputs, compete for GPU memory, and wait when the system cannot run them immediately.

Useful capacity should not be communicated as one tokens-per-second figure. It is better expressed as a conditional commitment: for defined input and output scenarios, a given concurrency maintains targets for time to first token, total response time, and rejection or waiting rate. This wording makes an internal promise testable against a repeatable benchmark and prevents extrapolation from a favorable demo.

Generative-model serving has separate metrics for queue time, time to first token, generation time, latency between tokens, and end-to-end latency. That distinction matters because two configurations with the same average throughput can provide very different experiences: one may begin responding promptly but finish slowly; another may delay the start because work has accumulated. Local-model guides and runtime comparisons should address these dimensions separately before recommending a configuration.

The operational question is not simply, “How many people are in the company?” It is, “How many active requests, of what size, and with what latency target must be sustained at the same time?” Thirty people with occasional use may imply low concurrency. By contrast, a few automations attaching long documents or producing lengthy outputs can saturate the same server. Measurement must represent those workloads, not an abstract idea of a user.

02

Define the service before selecting or expanding hardware

Start by describing expected demand in units the server can observe. Record the number of active requests, not only registered users; the distribution of input tokens; the allowed output maximum; task type; arrival rate; and availability and latency objectives. If historical traffic does not exist, state the assumptions explicitly and test conservative scenarios. An estimate does not become a fact merely because it is expressed with numerical precision.

Separate at least three workload classes. Short chat usually has small inputs and moderate outputs. A retrieval-augmented query can include document chunks and have a substantial input even when its answer is short. Drafting, extraction, or code generation can require long outputs and keep a conversation active for longer. Combining them in one average hides the cases that consume the most memory or block the queue.

For each class, define a budget: typical and maximum input tokens, typical and maximum output tokens, expected simultaneous requests, and experience limits. The latency objective should include at least time to first token and time to completion. You must also decide whether a user can cancel generation, what happens when a limit is reached, and whether priority classes exist. Without these rules, capacity depends on implicit decisions made by the runtime under pressure.

Comparing local models becomes useful only after this service contract is fixed. A smaller model may allow higher concurrency or more predictable responses on the same hardware; another may be justified by quality for a specific workload but require tighter limits. No configuration is universally sufficient: the decision must connect required quality to results measured under your own usage pattern.

Initial scenarios worth measuring separately

ScenarioInput and output to controlMain riskAcceptance indicators
Short chatShort input; capped outputQueueing caused by burstsTTFT and total time at defined percentiles
Retrieval queryLarge document input; short or medium outputPrefill work and KV-cache occupancyTTFT, KV-cache use, and waiting requests
Long-form generationMedium input; long outputProlonged memory retention and decodeTotal time, cancellations, and queue degradation
03

Build a memory budget, but do not mistake it for a guarantee

Memory available to serve a model is not limited to the published size of its weights. It must accommodate loaded weights, the key-value cache for active conversations, execution buffers and temporary workspace, runtime structures, and a reserve for variation and recovery. The exact distribution depends on the model, quantization, runtime, hardware, and configuration. A generic formula should therefore not be presented as if it were a universal measurement.

The KV cache is decisive for concurrency. It preserves the attention state needed to continue a sequence and grows with processed tokens. In a service, its occupancy changes by request: a long conversation, a large retrieved input, or a long output can retain a relevant share of memory much longer than a short question. Research on PagedAttention identifies management of this cache, including fragmentation and duplication under certain patterns, as a factor that constrains effective batch size.

It is reasonable to use a working budget for planning as long as it is labeled as an estimate. First, measure baseline memory with the model loaded and no traffic. Then observe how it changes while running each scenario with controlled lengths and increasing concurrency. Reserve capacity that is not assigned to nominal load. Finally, validate that the admission policy prevents the limit from being exceeded during a burst. The relevant result is observed behavior in the specific configuration, not an isolated calculation.

Runtime metrics can expose KV-cache utilization, running requests, and waiting requests, as well as prefill and decode measurements. These observations allow more cautious attribution of an incident: they do not by themselves prove one physical cause, but they help distinguish a growing queue from sustained cache pressure or slow generation. Preserve the configuration that produced every time series so analysis remains reproducible.

Process for estimating the memory budget

  1. 01Fix the model, available revision, quantization, runtime, driver, GPU, and context limits; record those values.
  2. 02Measure a baseline after loading the model and completing a warm-up with no test traffic.
  3. 03Run each scenario with one request and known input and output lengths; observe memory, KV cache, TTFT, and total time.
  4. 04Increase concurrency in small steps, keeping the scenario constant and recording queueing, errors, cancellations, and percentiles.
  5. 05Set an operating limit below the first point of instability, and verify that it leaves headroom for a burst or a late cancellation.
04

Prefill and decoding: two phases, two possible bottlenecks

A request does not consume resources in the same way throughout its lifetime. During prefill, the system processes the input to build the state generation will use. During decoding, it produces successive tokens and updates that state. A request with a long document can have a slow start even when its answer is brief; a long answer can begin quickly while retaining resources for a long time. Measuring only full duration erases that difference.

Time to first token is useful to the user and often reflects both queue waiting and the preliminary work needed to start generation. Latency between tokens, output-token time, and total time provide another view of the generation phase. Record actual input and output size as well, because variation in those lengths can explain apparent latency variation even when hardware has not changed.

Continuous batching can improve utilization by mixing work from different requests, but it does not remove memory limits or guarantee fairness across workloads. A broad-context workload can compete with short chats; scheduler decisions affect who starts first and who remains in the queue. A representative test should therefore include both homogeneous batches and a controlled mix of scenarios, reported separately.

Do not interpret a drop in average performance as an automatic diagnosis. It may result from arrivals exceeding service capacity, longer inputs, a high output limit, memory pressure, or scheduling policy. Instrumentation should provide enough context to identify correlations and, when the cause cannot be attributed, should state the uncertainty.

05

From VRAM to operating concurrency

Operating concurrency is the largest number of requests the service can admit for a defined scenario without missing its objectives. It should not be inferred directly from free memory or a theoretical maximum context. Memory may be sufficient while queueing or latency still exceeds the budget. Conversely, an acceptable result at one concurrency does not validate a workload with larger inputs or outputs.

Build a test matrix. On one dimension, define workload classes and their input and output lengths. On the other, increase simultaneous requests. For every cell, repeat the test after warm-up and collect percentiles for waiting, TTFT, generation latency, and end-to-end time. Record actually processed tokens, cache use, active and queued requests, cancellations, rejections, and errors. Averages may be added, but must not replace percentiles.

Capacity should be defined by the worst result you accept, not by the highest level that manages to complete one run. If p95 start time exceeds its target, the queue grows persistently, or memory errors appear, that concurrency is not operating capacity for that scenario. It can be retained as an investigated failure point for tuning limits, but not as a user-facing promise.

It is also important to test recovery after pressure. Following a burst, observe whether the queue returns to normal levels, whether memory is released as expected, and whether new requests regain usual latency. A system that completes a short test but does not recover promptly can be fragile for an internal API.

How to interpret a test-cell result

ObservationCautious interpretationInitial decision
TTFT is within target and the queue is stableThe scenario passes under this tested loadKeep it as a candidate and repeat
p95 TTFT rises; total time remains acceptableThe start experience is degradingReduce concurrency or limit input
The queue grows during the test windowArrivals may exceed service capacityApply admission, separate workloads, or expand capacity
Memory errors or unsolicited cancellationsThere is insufficient headroom for this loadLower limits and review memory reserve
06

Use explicit queues and admission before failure occurs

A queue is not necessarily an error: it can be a controlled decision that protects requests already in progress. The problem appears when it has no limit, when the user does not know they are waiting, or when the service continues accepting work it cannot complete within budget. An admission policy should decide, before resources are assigned, whether a request can enter, wait, receive a lower limit, or be rejected with a clear response.

Common controls include limits per user or credential, maximum input tokens, maximum output tokens, a maximum number of active requests, maximum queue length, and maximum wait time. Cancellation must release work and memory in a verifiable way. If priority classes exist, document them: priority does not create capacity; it only distributes limited capacity differently.

Degradation must be explicit and compatible with the use case. For example, an interface may ask the user to reduce attached documents, apply a disclosed output cap, or postpone a non-interactive task. It is not appropriate to silently truncate critical information or switch models without notice when that would alter the expected result. The policy should determine what happens before the runtime reaches an out-of-memory error.

Counters for waiting and running requests, together with metrics for prefill and decode work, help assess whether rules protect the service. Deliberately test a load above the admitted level: verify that new work is limited and that the latency of in-progress requests does not degrade uncontrollably. This overload test is as important as the nominal test.

07

A reproducible protocol on your own hardware

A useful test must be repeatable. Freeze and record the available model identity, its quantization, runtime, driver version, GPU type and count, context and output limits, batching parameters, and admission policy. If any of these variables changes, treat the result as a new measurement, not as an automatic continuation of the previous one.

Prepare a synthetic load based on the defined scenarios, without using real conversations unless there is specific authorization and appropriate controls. The load should set or record input and output lengths. Run warm-up separately from measurement, make several repetitions, and keep an observation period long enough to detect growing queues. Report p50, p95, and p99 when the sample volume supports their interpretation; alongside them, state the request count and test interval.

vLLM benchmarking documentation includes metrics such as TTFT, output-token time, and inter-token latency, and warns that prefix caching can inflate results when it is not controlled. If you use prefix caching in production, test it representatively and state whether prompts repeat; if it does not represent expected load, disable it or separate it into another scenario. A result dependent on unrealistic reuse is not a prudent capacity estimate.

Keeping only a dashboard is not enough. Save a configuration summary, load generator, parameters, aggregated results, and acceptance criteria. Reproducibility makes it possible to compare a model or runtime update and detect regressions. It also prevents a purchasing or deployment decision from depending on memories of a previous demonstration.

Recommended test sequence

  1. 01Set TTFT, total-time, waiting-rate, rejection-rate, and recovery objectives for each scenario.
  2. 02Warm up the service and exclude that phase from measured results.
  3. 03Test every scenario in isolation at increasing concurrency; record per-request results.
  4. 04Test a mixture of scenarios with a declared arrival pattern and compare results by class.
  5. 05Subject the service to a burst above its limit; validate admission, cancellation, and recovery.
  6. 06Repeat with the same configuration and document variation, failures, and changes from the initial hypothesis.
08

Decide from evidence: limits, model changes, or more hardware

With the results assembled, decide against the requirement rather than intuition. Keep the configuration if it meets priority scenarios with headroom and recovers after peaks. Reduce context or the output maximum when the product can do so explicitly and tests show that this restores the objectives. Limit concurrency if the demand pattern allows controlled waiting. Separating workloads may be preferable when long-document tasks interfere with interactive chat.

Adding GPUs or changing architecture is reasonable only after identifying the constraint you intend to relieve. If cache pressure dominates, the memory budget and context distribution are central. If long-input prefill misses TTFT, evaluate that phase separately. If model quality does not satisfy the task, more concurrency does not solve the problem. Evidence from one test does not automatically establish which alternative will be optimal without measuring it.

Changing models also requires repeating the matrix. Two models with similarly sized labels can use different context configurations, quantizations, and runtimes. The change can affect both baseline memory and generation behavior. In an internal comparison, communicate the full conditions and do not assign causality to weight size that has not been isolated.

If no configuration meets the requirement with acceptable limits, the honest decision may be not to deploy the shared service yet. A limited pilot with clearly defined scope is preferable to promising a general-purpose private assistant without a capacity budget. Local-model inventories and comparisons should present that conclusion as an operational possibility, not as a failure.

Decisions by observed bottleneck

Test findingChange to evaluateRequired validation
Long input misses TTFTReduce context, improve retrieval, or separate that workloadRepeat prefill with a representative distribution
Long output degrades everything elseCap output, cancel, or isolate long tasksMeasure chat queueing and latency during a mixed workload
Memory has no headroomLower concurrency or context; change hardware capacityBurst and recovery test
Quality is insufficient at sustainable limitsChange model or redesign the taskQuality evaluation and a new capacity matrix
09

Privacy, telemetry, and deployment checklist

Capacity instrumentation must not create a parallel repository of conversations. Generative-AI telemetry specifications warn that input and output messages, instructions, and arguments can contain sensitive information or personal data. Measuring concurrency usually does not require storing full text: token lengths, timestamps, pseudonymized identifiers, admission outcomes, and latency metrics are generally sufficient.

Before enabling detailed traces, define which fields are collected, for what diagnostic purpose, who can access them, how long they are retained, and how they are deleted. If prompts or responses are retained for debugging, the exception must have a justification, access controls, and a limited retention period. Also review whether seemingly harmless attributes, such as tool names or arguments, can reveal business information.

Minimum evidence for an internal deployment includes the service contract, scenarios, technical environment, load method, percentiles by scenario, behavior above the limit, admission policy, and telemetry treatment. Always distinguish what was measured, what was estimated, and what has not yet been tested. This discipline makes it possible to adjust the service without turning a demonstration figure into an unsupported guarantee.

Capacity is not permanent. Changes to the model, quantization, driver, runtime, context parameters, caching, or scheduling policy can invalidate previous results. Schedule a repeat of the test when material elements change, and monitor in production that actual input, output, and waiting distributions do not drift from approved scenarios.

Checklist before announcing internal capacity

  1. 01Does the service declare scenarios, input and output lengths, concurrency, and percentile targets?
  2. 02Were short chat, broad context, and long output separated, with results by class?
  3. 03Were queueing, TTFT, generation, total time, KV-cache use, rejections, and cancellations recorded?
  4. 04Was overload tested, and was admission verified to protect in-progress requests?
  5. 05Does the full technical configuration allow the test to be repeated?
  6. 06Does telemetry avoid conversation text except for a justified, protected, temporary exception?
  7. 07Were headroom, uncertainties, and conditions requiring a repeat test documented?

Open questions

  • No universal formula based only on VRAM or weight size can determine concurrency across all models and runtimes.
  • Acceptable thresholds for p50, p95, p99, waiting, and rejection depend on the product and are not set by the supplied sources.
  • Precisely attributing a performance decline can require additional instrumentation; metrics reveal correlations but do not always prove a single cause.
  • The effect of prefix caching, batching, and scheduling depends on configuration and actual prompt repetition, and must be measured in the local environment.
  • Synthetic-load results may not represent production traffic if input, output, or arrival distributions change.
10

Keep exploring

10

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction