Why does vLLM run out of memory before it serves a single request?
Few things are more frustrating than deploying an LLM serving container, watching the weights load into VRAM, and then immediately crashing with a CUDA out-of-memory error before handling a single request. When you inspect the traceback, no inference traffic has hit the server. This happens because vLLM operates on a fundamentally different memory allocation model than traditional PyTorch services.
In standard deep learning serving stacks, memory is allocated dynamically as requests enter the pipeline. If a batch is too large or sequences exceed expected lengths, the engine crashes under load. vLLM avoids dynamic runtime memory allocation by design. At engine initialization, it profiles the GPU, loads the model weights, measures intermediate activation requirements, captures CUDA graphs, and then aggressively allocates all remaining available VRAM into a static pool of PagedAttention key-value (KV) cache memory blocks.
Understanding this architecture makes the core diagnostic rule clear: an out-of-memory error at engine startup is a reservation problem, whereas an out-of-memory error during traffic is an admission problem. When vLLM fails to boot, the sum of the model weights, runtime profiling overhead, CUDA context, and minimum requested KV cache blocks exceeds the slice of VRAM allocated by your configuration.
- Weight loading: The engine loads unquantized or quantized model weights directly into device VRAM.
- Profiling pass: vLLM executes a dummy forward pass to measure peak activation memory and allocate workspace memory for CUDA graph execution.
- KV cache pool reservation: The engine evaluates remaining VRAM within the utilization ceiling and pre-allocates contiguous block tables.
- Block sufficiency verification: If the remaining VRAM cannot accommodate the minimum required blocks for the configured context length, initialization halts immediately with an exception.
When tuning vLLM, you are not debugging a memory leak. You are balancing a deterministic reservation equation. Resolving the crash requires adjusting the parameters that govern this initial reservation, such as tensor parallelism, quantization, context length and batch size, rather than defaulting to overprovisioning a larger GPU node.
What does gpu_memory_utilization actually reserve, and what is left for everything else?
The primary knob controlling vLLM memory reservation is the --gpu-memory-utilization engine argument. It defines the fraction of GPU memory the model executor is permitted to occupy, ranges from 0 to 1, and defaults to 0.92 in current releases (earlier versions defaulted to 0.9). It is a per-instance limit, so two vLLM instances sharing one card each need their own budget: the docs suggest setting 0.5 for each of two instances on the same GPU.
Many engineers misinterpret this setting as a dedicated quota for the KV cache. In reality, gpu_memory_utilization represents the ceiling for the entire vLLM model executor process. The memory consumed under this ceiling is partitioned into three distinct segments: the static model weights, the temporary activation buffers and CUDA graph capture memory, and the remainder, which forms the PagedAttention KV cache pool.
| Allocation Component | 80GB GPU at 0.90 Utilization | 80GB GPU at 0.95 Utilization | Role in vLLM Runtime |
|---|---|---|---|
| Model Weights (e.g. 70B FP8) | 35.0 GiB | 35.0 GiB | Static footprint loaded at engine boot |
| Activation & CUDA Graphs | 4.5 GiB | 4.5 GiB | Profiling workspace for graph capture |
| KV Cache Block Pool | 32.5 GiB | 36.5 GiB | PagedAttention dynamic token allocation |
| Reserved Headroom (Outside vLLM) | 8.0 GiB (10%) | 4.0 GiB (5%) | CUDA context, PyTorch allocator, NCCL buffers |
The headroom left outside the utilization ceiling is critical for system stability. This unreserved space houses the PyTorch runtime context, CUDA driver runtime states, NCCL communication buffers for distributed tensor parallelism, and external system processes. If you push gpu_memory_utilization close to 1.0, the PyTorch allocator or NCCL will run out of addressable physical memory during driver initialization or collective communication calls, triggering an unrecoverable CUDA driver crash.
As an operational rule of thumb, maintain a strict ceiling of 0.95 for gpu_memory_utilization on dedicated single-GPU nodes. In multi-GPU clusters where NCCL buffer allocations scale with tensor-parallel world sizes, keeping utilization at 0.90 or 0.92 ensures sufficient headroom for collective operations.
How does max-model-len set the size of the KV cache?
A common misconception when configuring vLLM is that reducing --max-model-len shrinks the total GPU memory footprint of the container. It does not. The docs define gpu_memory_utilization as the fraction of GPU memory to be used for the model executor, applied per vLLM instance, and the KV cache pool fills whatever is left under that ceiling. Declaring a shorter context therefore changes how that reserved space is divided per sequence, not how much of the card vLLM claims.
What max-model-len actually governs is the reservation granularity per sequence. In PagedAttention, the engine allocates physical memory in fixed-size blocks (typically 16 or 32 tokens per block). When a request arrives, the engine must ensure it can schedule tokens up to the maximum potential sequence length without running out of slots. Understanding KV cache memory calculation is essential for predicting this boundary.
- Per-sequence worst-case requirement: A larger context length increases the number of blocks an individual sequence might consume during generation.
- Maximum concurrent concurrency: With a fixed total KV cache pool (for example, 32 GiB), reducing max-model-len from 128k to 8k lowers the reservation ceiling per request, allowing the engine to schedule significantly more concurrent streams.
- Startup block threshold: vLLM requires sufficient memory to support at least one full-length sequence at startup. If the available KV cache pool cannot satisfy this single-sequence requirement, the engine refuses to start.
Modern open-weight models frequently ship with very long default context windows. If you load a large model on a single GPU without overriding --max-model-len, vLLM derives the context length from the model config and must fit cache blocks for that full window, which is why the docs list limiting the context length as a first-line way to reduce memory usage.
If your application pipeline only handles short customer support queries or summarization tasks, setting --max-model-len to that working ceiling instead of the model default immediately relieves startup memory pressure and maximizes the available concurrency slots within your existing context length requirements.
"No available memory for the cache blocks" - what is vLLM telling you?
When vLLM throws ValueError: No available memory for the cache blocks, it is reporting a definitive mathematical failure in its initialization formula. Specifically, after loading the model weights into VRAM and executing the memory profiling step for activations and CUDA graphs, the amount of memory remaining under the gpu_memory_utilization threshold is zero or insufficient to allocate even the baseline minimum of PagedAttention cache blocks.
To eliminate this error without purchasing larger hardware, you must reduce the static memory overhead consumed before the cache allocation step occurs. You have four technical levers:
- Weight quantization: If you are running 16-bit floating-point weights (FP16/BF16), quantize the model to FP8, AWQ, or GPTQ. Compressing a 70B model from BF16 (~140 GB) to FP8 (~70 GB) cuts static weight consumption in half, instantly opening tens of gigabytes for cache blocks.
- Tensor parallelism: Distribute model weights across multiple GPUs using --tensor-parallel-size. Sharding an unquantized model across 2, 4, or 8 devices divides the per-GPU static weight footprint proportionally.
- Eager mode execution: Disable CUDA graph capture by adding --enforce-eager. While this slightly increases inter-token latency during inference, it eliminates the dedicated workspace memory vLLM pre-reserves for CUDA graph execution trees.
- Context clamping: Restrict --max-model-len to the actual ceiling required by your workload to lower the minimum block allocation barrier.
Applying these adjustments directly targets the static side of the memory ledger. Tensor parallelism, quantization, limiting context length and batch size, and reducing or disabling CUDA graph capture are all listed in the vLLM documentation's guidance for machines that run out of memory. If you are diagnosing broader infrastructure issues, reviewing methods for resolving CUDA out of memory errors provides helpful context across training and serving environments.
Startup OOM or runtime OOM: which knob do you turn first?
When debugging vLLM memory failures, your remediation sequence depends entirely on whether the failure happens at boot or under production traffic. Turning the wrong configuration knob can unnecessarily degrade latency or throughput without addressing the root constraint.
If you experience a startup OOM, your focus must be on shrinking the static reservation. Follow this ordered triage process:
- Check model parameter size versus available VRAM. If unquantized model weights leave little room for cache blocks on a single card, either enable tensor parallelism or apply model weight quantization, both of which the vLLM docs recommend for memory-constrained deployments.
- Explicitly clamp --max-model-len to match your actual application payload instead of accepting the model default context window.
- Increment --gpu-memory-utilization from the default up to roughly 0.95 if physical GPU memory has available overhead.
- Add --enforce-eager to disable CUDA graph memory overhead during initial testing.
If the server boots cleanly but fails under load, you are encountering a runtime admission and preemption issue. When KV cache space is insufficient for all batched requests, vLLM preempts requests and logs a warning naming PreemptionMode.RECOMPUTE, recommending that you increase gpu_memory_utilization or tensor_parallel_size, or decrease max_num_seqs or max_num_batched_tokens.
| Failure Mode | Root Cause | Primary Knob to Turn | Secondary Optimization |
|---|---|---|---|
| Startup: CUDA OOM / No cache blocks | Static weights and graph buffers exceed utilization ceiling | --tensor-parallel-size or weight quantization | Set --max-model-len to workload ceiling |
| Startup: Driver / NCCL Init Crash | Zero headroom left for CUDA driver and communication | Reduce --gpu-memory-utilization to 0.90 or 0.92 | Verify no competing background processes |
| Runtime: High Preemption Rate | KV cache pool exhausted by high concurrency or long sequences | Switch to --kv-cache-dtype fp8 | Tune --max-num-seqs and enable chunked prefill |
| Runtime: Latency Spikes under Load | Memory bandwidth bottlenecked by massive 16-bit KV cache | Quantize KV cache to FP8 precision | Optimize batch sizing via max-num-batched-tokens |
For runtime memory pressure, quantizing the KV cache itself is one of the most effective optimizations. vLLM's cache_dtype option defaults to "auto", which stores the KV cache in the model data type, and CUDA 11.8 and later also accept fp8 (fp8_e4m3) and fp8_e5m2. Passing --kv-cache-dtype fp8 therefore stores each cached element in eight bits instead of the model's 16-bit dtype, cutting the VRAM each token occupies in the cache block pool. The benefit grows with sequence length, since longer contexts spend proportionally more memory on cache blocks, so you gain batch concurrency on identical hardware.
How do you verify the fix without waiting for it to fail in production again?
Do not rely on trial-and-error deployments to verify vLLM memory configurations. The vLLM initialization output provides explicit telemetry regarding memory allocation that you can inspect before sending production traffic.
During boot, vLLM prints a structured log summarizing memory profiling results. Locate the lines detailing available GPU blocks, KV cache size in GiB, and maximum concurrency. A healthy configuration will show a substantial block count (typically several thousand blocks) and clear headroom for parallel requests.
For deterministic, repeatable deployments in Kubernetes or container fleets, avoid relying entirely on dynamic percentage calculations. vLLM logs the exact --kv-cache-memory value that reproduces the current allocation, and passing it back on the next boot skips the memory-profiling measurement and the CUDA graph memory estimation pass, so startup is faster and every replica allocates identically. The docs caution that the value is only valid on the same GPU with the same initial free memory, and that a conservative value caps concurrency while an optimistic one fails at allocation time.
Once configured, validate your deployment with synthetic benchmarking tools such as the vLLM benchmark_serving script. Simulate worst-case concurrency and maximum prompt lengths to ensure the scheduler maintains zero preemption events under peak load before routing customer traffic to production model serving endpoints.
When does handing the engine to a managed endpoint cost less than tuning it?
Optimizing vLLM for high-throughput production workloads requires continuous operational effort. Between profiling memory utilization, balancing tensor parallelism across nodes, tuning KV cache precision, and managing preemption thresholds, infrastructure engineering time accumulates rapidly. When teams miss infrastructure forecasts or hoard expensive GPU instances out of fear of OOM crashes, compute budgets erode quickly.
The alternative is to move the operational boundary. If your team wants to consume high-performance open-weight models without managing GPU memory reservations, CUDA drivers, or vLLM configuration flags, a serverless inference endpoint provides an OpenAI-compatible API billed strictly per token with no egress fees. Where European data location matters, specific models such as Qwen3-Coder-30B-A3B and Hermes-4-405B are served from EU data centres (eu-north1). Reviewing how serverless inference billing works shows how eliminating idle GPU standby time reduces total compute spend.
| Deployment Model | Infrastructure Management | Pricing Mechanism |
|---|---|---|
| Serverless inference | Zero container or engine management; fully managed scaling | Per-token metering, no egress fees |
| Dedicated inference | Dedicated GPU hardware with managed runtime | Per GPU per hour |
| On-demand GPU VM | Full root access and custom vLLM / CUDA engine configuration | Per-second billing, no egress fees |
If your architecture requires custom engine binaries, private model adapters, or specialized CUDA kernels, managing the serving stack directly remains the right choice. For those workloads, Lyceum Dedicated Inference and On-demand GPU VM instances offer high-performance hardware with per-second billing and predictable provisioning, giving you complete control over your memory allocation equations.