
GPUSmith Article
vLLM KV Cache Capacity Planning: Context and Concurrency
Summary
- 01Plan from memory actually assigned to the KV cache. Account for model weights, runtime allocations, graph pools and workspaces before converting cache bytes into resident-token capacity.
- 02For ordinary full-attention layers, storage depends on cache-bearing layers, KV heads, head dimension and stored element width. Query-head counts and weight quantization are insufficient substitutes.
- 03Reserve prompt plus bounded output tokens for each request, round to cache blocks and retain an explicit local margin. Keep active-request limits and waiting-work limits separate.
- 04The article's hypothetical request envelopes describe storage feasibility. Reconcile them with startup allocation, then validate memory behavior, preemption, output quality and service latency under the intended workload.
- 05Treat cache reuse and nonstandard layouts as separate scenarios. Measure cold and warm conditions independently, verify local rank storage and let the limiting rank or cache pool determine capacity.
Inside this article
- 01Executive Summary
- 02Introduction and Background
- 03Key Changes in the Capacity-Planning Model
- 04From Model Architecture to Concurrent-Token Capacity
- 05Implementation Considerations and Process Changes
- 06Preemption Monitoring and the Admission Gate
- 07Parallelism and Nonstandard Cache Layouts
- 08Data Analysis and Evidence
- 09Case Studies and Real-World Examples
- 10Implications and Future Directions
- 11Frequently Asked Questions (FAQs)
- 12Conclusion
Executive Summary
vLLM KV cache capacity planning starts with the memory actually assigned to the key-value (KV) cache, rather than the accelerator's advertised capacity. For ordinary full-attention layers, architecture-derived storage per cached token is 2 × layers × KV heads × head dimension × bytes per element. The factors represent keys and values, all cache-bearing layers, and their stored scalar elements. Hugging Face publishes the corresponding half-precision calculation. [1] Multi-query attention (MQA) uses a single KV head, while grouped-query attention (GQA) uses an intermediate count between MQA and multi-head attention (MHA). [2] [3]
The worked scenarios below are hypothetical calculations, not benchmarks. With 32 layers, eight KV heads, head dimension 128, and two-byte cache elements, storage is 128 kibibytes (KiB) per token. An assumed 24 gibibytes (GiB) cache holds 196,608 tokens before an analyst-selected 10% reserve and whole-block rounding. A request budget of 8,192 prompt tokens plus 2,048 output tokens then supports a conservative envelope of 17 simultaneous requests, versus 34 with one-byte, 8-bit floating-point (FP8) storage under otherwise identical assumptions. The architecture equation supplies the calculation, and binary units use the National Institute of Standards and Technology (NIST) definition of GiB. [4] [5] [6] These counts describe storage feasibility; they do not predict latency or throughput.
The decisive reconciliation is between the calculated envelope and vLLM's startup report. The documented GPU KV cache size is a simultaneous token capacity; its Maximum concurrency estimate uses the configured maximum sequence length. [7] That estimate is separate from a scheduler sequence cap and an upstream admission limit. Ray Serve, for example, independently bounds ongoing requests per replica and queued requests per deployment handle. [8] [9] A practical admission gate should reserve prompt plus bounded output tokens, round to cache blocks, and retain an explicit local reserve.
Documentation observed on October 5, 2026 supports a canary process that changes one cache or scheduling control at a time. vLLM documents recomputation after cache-driven preemption. [10] Memory profiling must also account for allocator reservation and graph capture; PyTorch distinguishes allocated tensor memory from reserved allocator memory. [11] LMCache's calculator provides a useful tokens-to-memory and memory-to-tokens cross-check, but its answer should be reconciled to the deployment's actual graphics processing unit (GPU) allocation. [12] Promote a canary only when its measured memory behavior, preemption deltas, and service-level objective (SLO) results meet locally defined acceptance criteria. Aggregate latency histograms correctly across replicas rather than averaging independent percentiles. [13]
Introduction and Background
The admission-control question is specific: how many long-context requests can remain active on a self-hosted model before their cached history exhausts the remaining device memory? KV caching stores previous calculations for reuse during generation. [14] It therefore creates a memory obligation that grows with retained history, alongside the model's other allocations. A model that loads successfully has passed an initialization check; the deployment still needs a workload-specific capacity decision.
This report separates model weights, runtime allocations, and resident KV state. A useful planning identity is available executor memory minus measured non-cache memory equals the automatic cache budget. The current vLLM GPU worker implements that allocation path using profiled non-KV memory and an applied CUDA-graph memory estimate. [15] The identity is a bookkeeping framework, not a universal percentage allowance. PyTorch's memory profiler also cannot see allocations made directly through CUDA application programming interfaces (APIs), including some communication-library allocations. [16]
GPU Smith's first-party description frames cluster sizing around workload models and throughput targets. [17] In this report, that adjacent engineering perspective means the outputs are a reproducible worksheet, an admission envelope, and a validation plan. No consulting delivery, customer outcome, or measured serving performance is inferred from that description.
LMCache explicitly offers a calculator for the memory needed to cache a specified number of tokens. [18] Its utility is transparency at the token-storage boundary. The operator must still identify what memory budget is being entered, which cache layout is assumed, and whether stored state is GPU-resident or reusable from another tier.
The final decision has separate queue and device dimensions. Ray Serve exposes a per-handle queue limit in addition to its replica concurrency limit. [9] A platform should therefore document active request capacity, waiting capacity, and request-length policy separately. The following sections turn those distinctions into formulas and canary evidence, with version-sensitive behavior dated to the documentation actually inspected.
Key Changes in the Capacity-Planning Model
Replace advertised VRAM with a measured cache budget
Video random-access memory (VRAM) is not synonymous with cache bytes. NVIDIA advertises H200 with 141 GB of high-bandwidth memory and 4.8 TB/s memory bandwidth. [19] Those specifications describe the device, not the memory remaining after a particular model and execution stack initialize. They cannot be substituted directly into a resident-token calculator.
Use a measured allocation boundary. PyTorch provides different measures for tensor allocations and the caching allocator's total reservation. [11] Its CUDA documentation also describes private pools for graph capture. [20] NVIDIA's graph guidance explains that peak capture memory determines the private pool reservation, including temporary tensors. [21] Consequently, a steady-state tensor snapshot alone is an incomplete estimate of the runtime footprint.
The planning record should make the following distinctions explicit:
- Weights: record the loaded representation and distribution across ranks.
- Activations: profile the intended prefill and batching regime.
- Graphs: record the capture configuration and retained memory pools.
- Workspaces: include library and communication allocations in the observed footprint.
- KV pool: use the resulting cache allocation as the worksheet input.
- Reserve: document the operator's remaining margin separately from runtime overhead.
These categories are motivated by documented allocator, graph, and workspace behavior; PyTorch describes retained cuBLAS workspaces per handle-and-stream combination. [22] The worksheet should avoid counting the same reservation twice when combining process-level and allocator-level measurements.
Separate allocation, request limits, and reuse
A manual cache allocation is a distinct control. The current vLLM CLI says that a supplied kv_cache_memory_bytes ignores gpu_memory_utilization. [23] Treating both as independently additive budgets would therefore produce a misleading capacity estimate. Record which allocation path the deployment uses before interpreting a flag change.
The request's logical history and its physical cache allocation also differ. Transformers' DynamicCache grows as generation progresses, whereas StaticCache preallocates a configured maximum size. [24] [25] These are explanations of cache behavior in Transformers, not guarantees about vLLM's allocator. They illustrate why a calculator must state its allocation model rather than assuming every serving stack reserves the same way.
Prefix reuse adds another distinction. vLLM's documented prefix-cache design caches full blocks, while LMCache describes storing reusable state in CPU memory. [26] [27] Neither statement establishes a universal multiplier for concurrent GPU-resident requests. The conservative baseline should assume independent histories, with reuse credited only in a separately measured workload scenario.
From Model Architecture to Concurrent-Token Capacity
Derive storage per token
For uniform, conventional full-attention layers, define L as cache-bearing layers, Hkv as KV heads, D as head dimension, and s as bytes per stored element. Then bytes_per_token = 2 × L × Hkv × D × s. The leading factor accounts for keys and values. Hugging Face's published half-precision expression uses the same dimensions. [1]
Query heads are not KV heads. MQA shares keys and values across attention heads, as described in the originating paper. [28] GQA instead uses multiple KV heads, fewer than query heads. [3] Substituting the query-head count into a GQA calculation can materially overstate cache storage. A model's parameter count or family name is an insufficient input to this equation.
Read the deployed model configuration and retain its revision. As a concrete cross-check, Qwen's official Qwen2.5-7B-Instruct card lists 28 query heads and four KV heads. [29] The research also fetched its public configuration. This is an example of checking the architecture, not a claim that this model matches the hypothetical scenarios below.
The input checklist is:
- Layer count: include layers that actually allocate the modeled cache.
- KV heads: use the key-value count, not the query count.
- Head dimension: read the explicit field or verify the architecture's derivation.
- Cache dtype: use the stored cache type, independently of weight quantization.
- Layer exceptions: identify sliding-window, cross-attention, latent, or hybrid state.
- Model revision: preserve the configuration used for the calculation.
For a two-byte scalar, bfloat16 (BF16)'s documented representation is 16 bits. [5] NVIDIA describes FP8 E4M3 with a sign bit, four exponent bits, and three mantissa bits, totaling eight bits. [6] These representation widths explain the theoretical storage ratio, but actual cache allocation can also include scales, alignment, and implementation-specific state.
Convert bytes into admission reservations
Use exact byte units. NIST defines 1 GiB as 1,073,741,824 bytes, distinct from decimal GB. [4] If the supplied cache budget is B bytes and block size is b tokens, the homogeneous estimate is T = b × floor(B / (b × bytes_per_token). LMCache's reverse calculator addresses the same memory-to-token question. [12]
The recommended worksheet then computes usable_tokens = b × floor((1 - reserve_fraction) × T / b). For request i, reserve b × ceil((prompt_tokens_i + output_budget_i) / b). Admit it only when the sum of existing reservations plus this new reservation fits the usable budget. These are analyst-defined conservative rules, not a documented vLLM scheduling algorithm.
Prompt and output percentile budgets are useful planning inputs, but percentile sums are not automatically a joint percentile. Preserve a hard request-length cap or calculate the distribution of combined prompt and output demand. Build separate workload classes when long-document requests and short interactive requests have materially different reservations. Prometheus likewise cautions that independently calculated quantiles cannot be meaningfully averaged. [30]
For mixed cache layouts, replace a single global coefficient with verified per-layer or per-group storage. Transformers documents that sliding-window and chunked layers stop growing at their window or chunk limit. [31] A linear full-context equation applied indiscriminately to those layers is therefore a scenario assumption, not an exact architecture model.
- 01Establish the cache budget
Use the measured allocation boundary. For automatic allocation, subtract measured non-cache memory from available executor memory.
- 02Derive storage per token
For ordinary full-attention layers, use cache-bearing layers, KV heads, head dimension and bytes per stored element.
- 03Round capacity and retain reserve
Convert cache bytes to whole-block token capacity. Apply the locally selected reserve and round the usable budget to whole blocks.
- 04Reserve each request's demand
Round prompt plus bounded output demand to blocks. Admit a request only when existing reservations plus its reservation fit the usable budget.
- 05Reconcile startup capacity
Compare startup cache tokens with the worksheet. Investigate units, rank layout, block rounding and cache groups before changing admission limits.
- 06Validate the canary
Promote only when measured memory behavior, preemption deltas and SLO results satisfy locally defined acceptance criteria.
Promote when the recorded workload meets local memory, quality and latency criteria.
Investigate discrepancies between byte-derived capacity and startup tokens before changing admission limits.
Only the deployment's completed evidence record can turn that estimate into an operational concurrency commitment.
Implementation Considerations and Process Changes
Treat controls as distinct canary hypotheses
Table 1 summarizes control meanings in documentation observed on October 5, 2026. The CLI defines model length as prompt plus output, documents manual cache-budget precedence, and exposes the scheduling controls; the quantized-cache documentation provides the FP8 compatibility caveats. [32] [23] [33]
| Control | Capacity question | Recommended canary evidence |
|---|---|---|
--gpu-memory-utilization | How much executor memory is available for automatic allocation? | Startup cache allocation, peak runtime memory, and headroom. |
--kv-cache-memory-bytes | Is the cache budget explicitly fixed per GPU? | Actual allocation and measured non-cache footprint. |
--max-model-len | What prompt-plus-output length is permitted? | Longest accepted request and startup length feasibility. |
--max-num-seqs | How many sequences may be scheduled per iteration? | Running population, preemption, and latency under load. |
--max-num-batched-tokens | How many tokens may be scheduled per iteration? | Prefill pressure, runtime memory, and latency distribution. |
--kv-cache-dtype | What storage representation and backend are used? | Allocation change, supported kernels, and output-quality comparison. |
The table is a set of experiment questions, not recommended universal defaults. In particular, a scheduler sequence cap is not an assurance that every permitted sequence can reach the maximum context simultaneously. The admission worksheet remains responsible for that storage obligation.
The current CLI's defaults can be resolved through engine configuration rather than a single universal value for all workloads. [32] Preserve the effective configuration printed by the exact deployed release. Avoid treating an old tuning article's flag defaults as a substitute for the current executable's help and startup output.
Reconcile allocation and logged concurrency
Capture the startup report before generating load. The documented maximum-concurrency estimate uses the configured maximum sequence length, while the token-capacity log describes simultaneous cache capacity. [7] For homogeneous full attention, compare logged tokens against byte-derived whole-block capacity. If they differ, investigate units, local rank layout, block rounding, and runtime cache grouping before changing admission limits.
Manual allocation still needs runtime validation. The current GPU worker skips the automatic cache-sizing profile when a manual byte budget is supplied, but still runs a profile step for compilation. [34] This is not evidence that the explicit budget will coexist safely with every later workload peak. The canary must exercise the intended prompt and batching regime.
Recommended deployment records include:
- Build identity: image digest, package version, and model revision.
- Effective flags: resolved settings, rather than remembered defaults.
- Allocation: bytes per rank and the startup token-capacity report.
- Workload: prompt and output distributions, arrival pattern, and cancellations.
- Runtime: device, driver, backend, graph mode, and cache representation.
- Acceptance: explicit memory, quality, and latency criteria.
Graph-capture memory can change the allocation boundary, so retain graph configuration beside the cache calculation. [21] Kubernetes GPU resource requests and limits, when both are set, must be equal; that pod resource assignment should not be interpreted as a token admission budget. [35]
Evaluate FP8 as a separate change
FP8 capacity gain is theoretical until measured. ROCm's versioned precision documentation notes encoding differences between MI300-native FP8 and NVIDIA H100 formats. [36] A dtype name alone is therefore not a complete compatibility specification. Verify the target device, attention backend, scale handling, and installed vLLM release together.
Scaling can add state: NVIDIA's Transformer Engine example uses a 32-bit floating-point (FP32) scale factor per tensor. [37] That is an illustration of low-precision representation overhead, not a statement of vLLM's exact scaling layout. Compare measured cache capacity and representative output quality before accepting a storage-based request-count increase.
Preemption Monitoring and the Admission Gate
Observe pressure and distinguish metric types
The current vLLM metrics documentation exposes vllm:kv_cache_usage_perc, running and waiting request gauges, and preemption metrics. Cache usage equal to 1 means full utilization. [38] It also documents a cumulative vllm:num_preemptions counter and a vllm:request_num_preemptions histogram for per-request preemptions. [39] Record the actual exported series on the deployed endpoint; registration names and exporter wire-format suffixes should not be assumed interchangeable.
A counter accumulates and may reset after restart, while a gauge can rise or fall. [40] [41] For the canary, calculate counter deltas or rates across an identified observation window, handle restarts, and retain denominators such as completed requests. A cumulative count without its window cannot establish whether a configuration change improved pressure.
The current vLLM logger separately describes statistics tracked over its local logging interval. [42] Keep those interval values distinct from lifetime cumulative series. Correlate cache occupancy, waiting population, preemption, and request latency rather than treating any single graph as the entire admission decision.
Use the following locally defined gate:
- Length gate: reject or reroute requests exceeding the configured length policy.
- Token gate: reserve block-rounded prompt and bounded output demand.
- Sequence gate: keep the running population within the scheduler's limit.
- Queue gate: cap waiting work independently of running work.
- SLO gate: retain the measured latency and quality criteria for promotion.
- Recovery gate: confirm restart and drain behavior before broad rollout.
For an outer serving layer, Ray documents separate per-replica ongoing and per-handle queued limits. [8] [9] Its queue bound includes the Serve proxy. [43] Translate the token gate into that layer deliberately; no numerical equivalence between a Ray request cap and vLLM cache capacity is established by those APIs.
Establish service-level evidence
Measure time to first token (TTFT), generation latency, throughput, and the waiting population for the same workload window. Choose the SLO thresholds locally; the sources inspected do not supply a universal latency acceptance threshold for the hypothetical deployment. Prometheus supports aggregating histograms before computing a service percentile. [13]
Do not average replica-specific percentiles to report service-wide latency. [30] Preserve warm-cache and cold-cache conditions separately, together with offered load and completion counts. A run with fewer admitted requests can reduce preemption while worsening queue delay, so promotion should evaluate both the device and the waiting path.
An out-of-memory (OOM) failure requires a different diagnostic from cache pressure. PyTorch distinguishes reserved allocator memory from live tensor memory and documents that releasing unused cached allocations does not free occupied tensor memory. [44] Revisit runtime peaks and allocation accounting rather than assuming every OOM is repaired by lowering a logical context cap.
Readiness is a traffic gate: Kubernetes does not send Service traffic to a pod whose containers report not ready. [45] Use it to hold a newly started replica out of rotation until its allocation and health checks complete. Autoscaling may adjust desired scale based on observed metrics, but it is a separate control loop. [46]
Parallelism and Nonstandard Cache Layouts
Calculate on the limiting rank
Tensor parallelism (TP) distributes model computation across devices; pipeline parallelism (PP) partitions layers. vLLM documents that sharding weights can leave more memory for KV cache and that these parallelism strategies introduce their own overheads. [47] The worksheet should therefore describe local cache-bearing layers and locally stored KV heads, rather than dividing a single-device answer by device count without checking the implementation.
For an evenly partitioned full-attention model (Hypothetical Example), local storage can be represented as 2 × local_layers × local_KV_heads × D × s. Equal head sharding and equal layer partitioning are explicit assumptions. Replicated heads, uneven stage assignments, or other layouts require the actual local storage counts. Calculate every rank and let the most constrained rank determine the admission envelope.
Similarly, multiple cache groups require group-aware reconciliation. The current vLLM cache utility rounds request requirements to pages and can use a separate host pool whose smaller concurrency limit applies. [48] Consequently, one token-count division may not reconstruct the runtime's estimate for a hybrid layout. Use the runtime report as evidence and retain the architecture worksheet as an explanation of the assumed components.
Bound the applicability of the linear equation
Mistral's original Mistral 7B description combines GQA with sliding-window attention and a rotating cache bounded by the window. [49] [50] This is a documented architecture example, not a general statement that every model bearing the same family name has identical cache growth.
Speculative decoding also changes the resource model. IBM explains that proposals may come from a smaller draft model or part of the main model itself. [51] Its implementation describes sharing cached history across candidates and notes memory competition with dynamic batching. [52] [53] Do not assume either a duplicate complete cache for every candidate or zero additional state. Reprofile the specific speculation configuration.
Multimodal serving can include additional encoder state. As a version-scoped example, AWS Neuron's vLLM-Neuron encoder-cache design stores vision encoder outputs in preallocated high-bandwidth memory and reserves scratch space. [54] [55] That evidence concerns Neuron, not CUDA-vLLM byte allocation, but it demonstrates why text-only KV accounting should not be presented as complete multimodal memory sizing.
The implementation checklist is:
- Sliding windows: sum full and bounded layers separately.
- Hybrid models: identify all runtime cache groups and state types.
- Latent layouts: replace the ordinary KV-head equation with verified storage dimensions.
- Speculation: include proposal state and the selected backend's allocations.
- Multimodal inputs: include encoder caches and supported modality limits.
- Parallel ranks: record local storage and the limiting pool.
The worksheet should mark unverified components explicitly. LMCache's embedded calculator exposes additional precision inputs for a specialized cache layout, reinforcing that its model selection is more than a generic parameter-count lookup. [56]
Data Analysis and Evidence
Transparent capacity worksheet (Hypothetical Example)
The following calculations isolate architecture and request demand. They assume 32 uniform full-attention layers, head dimension 128, 24 GiB of available KV cache per GPU, no sharing, and no TP or PP. Cache block size is an analyst-selected 16 tokens, with a 10% capacity reserve applied before another whole-block rounding step. These are scenario inputs, not vendor defaults or measured allocations. The equation is architecture-derived; dtype widths and GiB units come from the cited definitions. [1] [5] [6] [4]
Table 2 shows exact derived outputs. Prompt and output columns are reservations, not observed workload percentiles. Every number in the table is calculated from the stated hypothetical inputs.
| Architecture / cache | KV heads | Bytes per token | Raw tokens/GPU | Usable tokens/GPU | Prompt + output budget | Request envelope |
|---|---|---|---|---|---|---|
| MHA, two-byte | 32 | 524,288 | 49,152 | 44,224 | 8,192 + 2,048 | 4 |
| GQA, two-byte | 8 | 131,072 | 196,608 | 176,944 | 8,192 + 2,048 | 17 |
| MQA, two-byte | 1 | 16,384 | 1,572,864 | 1,415,568 | 8,192 + 2,048 | 138 |
| GQA, one-byte FP8 | 8 | 65,536 | 393,216 | 353,888 | 8,192 + 2,048 | 34 |
| GQA, two-byte, long | 8 | 131,072 | 196,608 | 176,944 | 32,768 + 8,192 | 4 |
| GQA, one-byte FP8, long | 8 | 65,536 | 393,216 | 353,888 | 32,768 + 8,192 | 8 |
| GQA, two-byte, extended | 8 | 131,072 | 196,608 | 176,944 | 65,536 + 8,192 | 2 |
The comparison holds cache allocation constant to expose the effect of retained history and KV heads. It does not hold compute cost or output quality constant across architectures. In particular, the MQA row is a storage comparison, not a recommendation to replace an existing model; the original MQA paper describes shared keys and values as its architectural change. [28]
A spreadsheet can reproduce the table with B = 24 × 2^30, bytes_per_token = 2 × 32 × Hkv × 128 × s, raw_tokens = 16 × FLOOR(B / (16 × bytes_per_token), and usable_tokens = 16 × FLOOR(0.90 × raw_tokens / 16). Then set request_tokens = 16 × CEILING((prompt + output) / 16) and request_envelope = FLOOR(usable_tokens / request_tokens). These are the report's formulas, using the documented binary-unit convention. [4]
For a mixed workload, the admission test is the sum of individual reservations, not the number of requests multiplied by an average observed length. Output budgets must remain bounded. LMCache's memory-to-token calculator is a useful independent interface check, but the inspected interface did not expose its JavaScript formula or complete unit assumptions. [12] Retain the worksheet's exact byte inputs instead of silently equating unlike calculator labels.
Canary record and promotion evidence
Table 3 is an unfilled evidence template, dated October 5, 2026. No rows contain measured results. Populate the startup columns using vLLM's reported cache tokens and maximum-concurrency estimate. [7]
| Canary change | Predicted capacity | Startup evidence | Preemption evidence | SLO and quality evidence |
|---|---|---|---|---|
| Baseline | Worksheet bytes/token and token envelope | Record bytes/rank, tokens, and context denominator | Record start/end counters and observation window | Record offered load, completions, and latency histograms |
| Allocation control | Recalculate using proposed cache bytes | Compare actual allocation with prediction | Compare normalized deltas with baseline | Check memory peaks and local SLO criteria |
| Sequence/batch control | Keep storage equation; revise running limits | Confirm whether allocation also changed | Record pressure under the same workload | Check both queue delay and generation latency |
| Cache dtype | Recalculate scalar width and known overhead | Measure actual token-capacity change | Check pressure at controlled offered load | Compare output quality and latency with baseline |
The template prevents predicted and measured capacity from being merged into one result. An outer replica limit remains a separate evidence field when Ray Serve is used. [8] For latency comparisons, aggregate histogram observations consistently rather than comparing mismatched percentile summaries. [13] A pass means the local criteria were met under the recorded workload, not that the setting is transferable to another deployment.
Case Studies and Real-World Examples
Ceph's cold and reusable-cache test
Ceph's engineering article provides a neutral example of controlling cache state. Its experiment contrasts computational prefill with a 100% remote-storage cache-hit condition across context lengths. [57] That comparison concerns TTFT and reused history; it does not establish a general GPU-resident concurrency increase.
The published procedure measures an initial computational run, restarts vLLM to clear GPU and CPU cache effects while retaining remote state, then measures retrieval from that state. Before the next context length, it stops serving and removes the remote cache blocks. [58] Those controls help distinguish a cold computation path from a reusable-storage path.
Ceph also publishes configuration and command-line material for reproducing its experiment. [59] The inspected page did not provide a complete immutable build identity for every component. Its results should therefore remain attributed to the published setup rather than treated as evidence for a different current release.
For capacity planning, adapt the methodological distinction. First establish independent-history resident capacity. Then measure a separate shared-prefix workload, preserving the cache state at the beginning and end of each test. LMCache's CPU-offloading documentation describes retaining and reusing cache from CPU memory; its connector can store and load KV state. [27] [60] That reusable tier has its own storage and transfer behavior.
The recommended test record includes:
- Cold state: specify which device, host, and remote caches were cleared.
- Warm state: specify exactly which histories remain reusable.
- Reuse pattern: preserve prefix identity and request ordering.
- Resident pressure: measure GPU allocation and running population separately.
- Transfer path: record the participating host or remote tier.
- Comparison window: use matching workload and latency accounting.
This procedure should prevent a warm-start latency result from being interpreted as a larger physical GPU cache. The relevant outcome is a documented operating mode with its own capacity and latency evidence. Ceph's explicit remote-cache clearing step is a useful model for preserving that distinction between repeated test runs. [58]
- Establish resident capacity with independent histories before measuring reuse.
- For mixed workloads, sum individual reservations and keep output budgets bounded.
- Measure shared-prefix workloads separately, preserving cache state at the beginning and end of each test.
- Ceph's comparison concerns time to first token and reused history; it does not establish a general resident-concurrency increase.
Credit reuse only in a separately measured workload scenario. Preserve cold and warm cache conditions independently.
Implications and Future Directions
Capacity planning should become a versioned operational artifact. Preserve the model configuration, calculation inputs, startup allocation, and canary record together. GPU Smith's description includes KV-cache and scheduler tuning within its serving-optimization scope. [61] The corresponding engineering implication is to make assumptions inspectable, rather than presenting a request count without the workload and runtime conditions that produce it.
The automation boundary should remain explicit. Kubernetes' HorizontalPodAutoscaler is a periodic control loop, with a documented default synchronization interval of 15 seconds. [62] Autoscaling evidence does not eliminate the need for a bounded queue and an immediate admission policy at existing replicas. Ray's per-handle queue semantics likewise require care when a service has several ingress paths. [43]
Low-precision and speculative execution should have separate refresh triggers. ROCm's precision tables distinguish encodings and device support. [36] IBM's description shows that speculative decoding can use different proposal structures and compete with batching for memory. [51] [53] A planner that stores only model name and nominal dtype will miss implementation details that matter to the next deployment.
Recommended refresh events are:
- Model change: repeat configuration extraction and architecture accounting.
- Runtime upgrade: recheck flags, metrics, cache groups, and allocation logs.
- Backend change: revalidate kernels, scales, and graph memory.
- Workload shift: revisit prompt and output demand distributions.
- Parallelism change: recalculate local rank storage and limiting capacity.
- Reuse change: rerun cold and warm evidence under controlled cache state.
The report does not establish a single ideal safety reserve or preemption threshold. Those remain local decisions tied to workload risk and SLOs. Service-wide latency must be calculated from appropriate aggregate observations; Prometheus documents why averaging quantiles is statistically unsound. [30] Keep that evidence beside capacity results so a storage improvement cannot silently substitute for a service improvement.
Frequently Asked Questions (FAQs)
What determines vLLM maximum concurrent requests?
Use the minimum of storage-feasible reservations, the scheduler's sequence allowance, the upstream admission limit, and the population that meets the local SLO. vLLM's startup concurrency estimate uses maximum sequence length, while Ray's ongoing limit is per replica. [7] [8] Neither alone establishes production concurrency.
Does reducing max-model-len automatically create more cache bytes?
Treat it as a request-length policy first. The CLI defines it as prompt plus output; the allocation path is controlled separately. [32] Reobserve startup allocation after changing it, since runtime profiling and execution configuration may also change. Distinguish a smaller request denominator from a larger physical cache pool.
Is a KV calculator a complete GPU-memory planner?
Its token-storage answer is a component estimate. LMCache documents memory-required and tokens-from-memory calculator modes. [18] The operator must provide an appropriate cache budget and separately account for weights, runtime pools, and other state. PyTorch's allocator documentation distinguishes live tensors from reserved memory. [11]
Can FP8 double the admitted request count?
Halving stored scalar width doubles idealized token capacity when every other assumption remains fixed. Actual promotion requires measured allocation, supported kernels, quality, and SLO evidence. vLLM documents format and backend conditions for quantized KV cache. [33] The hypothetical table's arithmetic is a planning estimate, not a performance guarantee.
Conclusion
The useful output of vLLM KV cache capacity planning is an admission rule with a traceable allocation boundary. Start with the deployed architecture, distinguish KV heads from query heads, use the stored cache representation, and retain every assumption that changes local rank storage. The formula explains ordinary full-attention state; the runtime evidence determines whether that explanation matches the actual deployment.
Convert measured cache bytes into block-rounded token capacity, then reserve prompt plus bounded output demand for admitted requests. A worksheet and a memory-to-token calculator can cross-check that boundary. LMCache explicitly supports the reverse calculation from available GPU random-access memory (RAM). [12] Keep safety margin as a documented local decision rather than presenting it as a vendor requirement.
Promotion should depend on a completed canary record. It should connect predicted capacity, startup allocation, pressure metrics, output quality, and correctly aggregated service latency. Prometheus' histogram guidance supports the aggregate latency evidence required for that comparison. [13] Separate the active population from waiting work and preserve the workload distribution used to establish the result.
Finally, retain cold-cache and reuse scenarios independently, identify the limiting rank or cache pool, and repeat validation when models or execution settings change. The report's numerical examples show how context and architecture reshape storage demand. Only the deployment's completed evidence record can turn that estimate into an operational concurrency commitment.
External Sources (62)
About
GPUSmith
Plan and build private AI compute with GPU Smith. We connect workload requirements with GPU hardware, networking, deployment and operating decisions so your infrastructure fits the work it must perform.
GPU Smith provides independent engineering and research for private AI infrastructure. We help technical buyers and operators reason about GPU workload sizing, hardware selection, procurement, deployment and inference operations, from the compute system to the networking, power and cooling requirements around it.
Start with the workload
Useful infrastructure decisions begin with the models, data, throughput, latency and operating constraints a team actually has. GPU Smith helps connect those requirements with system design and deployment choices. The goal is a privately controlled AI environment whose capacity and operational demands are understood before a purchasing decision.
Hardware and supplier research
Our public reference library covers GPUs, complete systems and networking components. The server-vendor directory and manufacturing research help teams investigate suppliers, compare options and follow sources behind technical claims. We distinguish manufacturer specifications from measured benchmarks and advertised capabilities from tested configurations. Reference pages are research resources, not a live inventory listing or a binding equipment quote.
Deployment and operations
GPU Smith's engineering scope includes infrastructure integration, networking, power, cooling and the practical operation of inference workloads. Our published articles explain the tradeoffs behind deployment, procurement and ongoing operations so teams can ask better questions and document decisions.
Work with GPU Smith
Explore the hardware library, GPU references, networking references and vendor research. Contact GPU Smith with your workloads and deployment constraints to discuss a project and confirm current scope, pricing and availability.
Inclusion of a manufacturer or product in our research does not imply a partnership, certification or endorsement.
Disclaimer
This document is provided for informational purposes only. No representations or warranties are made regarding the accuracy, completeness, or reliability of its contents. Any use of this information is at your own risk. GPUSmith shall not be liable for any damages arising from the use of this document. This content was generated with assistance from artificial intelligence tools, which may contain errors or inaccuracies. Readers should verify critical information independently. All product names, trademarks, and registered trademarks mentioned are property of their respective owners and are used for identification purposes only. Use of these names does not imply endorsement. This document does not constitute professional or legal advice. For specific guidance related to your needs, please consult qualified professionals.