LLM Inference: CPU and PCIe Bottlenecks Behind Low GPU Utilization
In LLM inference, low observed utilization can reflect low request volume, small batches, host overhead or other bottlenecks. Decode can also be limited by GPU memory bandwidth despite ongoing kernel activity. A generative model produces one output token per forward pass, so a request runs the model many times, and between passes the CPU has to schedule the next batch and launch its kernels. That launch work is a fixed cost per step, while the GPU's share of the work shrinks with the batch size: at small batch sizes the CPU overhead can exceed the GPU run time, and the GPU goes idle between kernel calls. PCIe adds a second wait, because peak bandwidth between host memory and the GPU is far lower than the GPU's own memory bandwidth, so every avoidable host-device copy in the serving loop widens the gaps.
You can confirm which of the two you have before changing anything. Profile a serving run with NVIDIA Nsight Systems or PyTorch Profiler and read the CPU and GPU rows together: short GPU bursts separated by gaps while a CPU thread stays busy point to launch and scheduling overhead, and long memory-copy spans point to host-device transfers. The fixes differ, so measure first.
- Batch more requests per step: high-throughput LLM serving depends on batching enough requests at a time, and the KV cache memory each request needs is what limits the batch size, so cache memory management directly decides how full the GPU runs.
- Use continuous batching: iteration-level scheduling runs the model one iteration at a time on the batch, so finished requests leave and newly arrived requests join without waiting for the whole batch to complete.
- Capture the decode step in a CUDA Graph: a graph replay submits the whole step's work with a single launch call and skips the Python, C++ and driver overhead of launching each kernel.
- Cut and batch host-device copies: keep tensors on the GPU between steps, use pinned host memory for the transfers that remain, and combine many small transfers into one to remove most of the per-transfer overhead.
If both CPU and GPU utilization are low during inference, check for low request volume, serialization, I/O waits or synchronization, so look at batching and the profiler timeline before you look at faster hardware.
The Data Loading Bottleneck: Feeding the Beast
One common cause of low GPU utilization is data starvation. If your GPU is waiting for the CPU to finish preprocessing or for the disk to read the next batch, your utilization will plummet. In enterprise clusters, I/O wait is a common contributor to idle GPU time, and it is rarely visible without profiling.
Diagnosing Data Pipeline Bottlenecks
To diagnose this, monitor your CPU utilization alongside your GPU. If your CPU is pegged while the GPU is low, data preparation is one possibility, but launch overhead or another CPU-bound stage can look similar. Confirm with a timeline before changing the loader. You can resolve this by optimizing your DataLoader configuration. In PyTorch, for instance, increasing the num_workers parameter allows for multi-process data loading, which parallelizes the fetching and transformation of data. However, setting this too high can lead to shared memory issues or CPU thrashing.
- Enable Pin Memory: Use pin_memory=True in your DataLoader. This enables faster data transfer from the CPU to the GPU by using page-locked memory, bypassing an extra copy step in the host memory.
- Prefetching: With num_workers above 0, each worker loads prefetch_factor batches in advance (2 by default); raise it if the next batch is still not ready when the current one finishes. This masks the latency of data preparation.
- Storage Throughput: Ensure your underlying storage can handle the IOPS required. Moving from SATA SSDs to NVMe storage or using high-performance parallel file systems is often necessary for H100 clusters.
Teams often overlook the impact of complex data augmentations. If you are performing heavy image, video or audio augmentations on the fly, consider precomputing or caching suitable transformations, or moving them to the GPU with NVIDIA DALI, which provides GPU-accelerated building blocks for loading and processing image, video and audio data. For text, tokenize ahead of time and cache the result. Measure the effect rather than assuming the GPU will remain saturated.
Batch Size and Tensor Core Saturation
GPU architecture is designed for massive parallelism. If your batch size is too small, you are not providing enough work to fill the thousands of CUDA cores available. This results in 'under-utilization' even if the GPU is technically active. For modern architectures like Hopper (H100), the Tensor Cores require specific alignment to reach peak TFLOPS.
Gradient Accumulation for Small GPUs
Choose a physical microbatch that fits with adequate memory headroom, then measure throughput and training quality. Gradient accumulation combines gradients across multiple microbatches to increase the effective batch per optimizer update. It does not increase the work in each microbatch, guarantee better GPU saturation or make a model fit when its weights and training state alone exceed device memory. Model-state limits may require sharding, offloading or different hardware.
- Shape alignment: Tensor Core efficiency depends on matrix dimensions, dtype, kernels and GPU generation. Suitable multiples can help, but there is no universal rule that batch sizes must always be powers of two or multiples of the warp size.
- Mixed precision: FP16 or BF16 can reduce tensor memory and accelerate supported kernels. Total memory is not automatically halved because optimizer state, master weights or other tensors may stay FP32. FP8 needs supported hardware and software plus numerical validation.
- Memory Profiling: Use tools like torch.cuda.memory_summary() to identify where your VRAM is going. Often, large activations or unoptimized model architectures are the real culprits behind small batch constraints.
The goal is to reach a state where the GPU is the bottleneck, not the data pipeline. If you increase the batch size and the utilization stays low, the issue likely lies in kernel launch overhead or synchronization delays.
Kernel Launch Overhead and CUDA Graphs
In deep learning, a single forward pass consists of hundreds or thousands of individual operations (kernels). Each time the CPU tells the GPU to run a kernel, there is a small amount of overhead. If your kernels are very small or execute very quickly, the time spent launching them can exceed the time spent actually running them. This is known as being 'CPU-bound' on the launch side.
This is a frequent issue with models that have many small layers or complex branching logic. To mitigate this, NVIDIA introduced CUDA Graphs. Instead of launching kernels one by one, a CUDA Graph allows you to 'record' a sequence of operations and launch them as a single unit. This drastically reduces the CPU-to-GPU communication overhead.
PyTorch torch.compile captures and optimizes execution graphs. TorchDynamo performs graph capture; compiler backends such as TorchInductor generate optimized code and can fuse suitable operations. Compilation may reduce kernel launches and memory traffic, but graph breaks, dynamic shapes and compilation overhead affect the result. Benchmark representative steady-state throughput and startup cost, and verify correctness.
Distributed Training and Communication Latency
When scaling across multiple GPUs or nodes, the bottleneck often shifts from local compute to network communication. In a distributed environment, data-parallel workers synchronize gradients. PyTorch DistributedDataParallel typically launches all-reduce operations as gradient buckets become ready during backpropagation, overlapping communication with compute where possible. If your network is slow or your synchronization strategy is inefficient, your GPUs will sit idle waiting for data from their peers.
Distributed training efficiency depends on the selected instance, topology and network fabric. InfiniBand or RoCE can help communication-heavy multi-node workloads, but the required bandwidth depends on model size, parallelism strategy and batch size. Measure scaling efficiency on the actual configuration before adding GPUs.
- Check communication: NCCL_DEBUG=INFO emits diagnostic logs. Use a profiler such as Nsight Systems for a timing trace, and run appropriate NCCL tests to measure collective performance; debug logs are not themselves a GPU timeline.
- Gradient bucketing: DDP overlaps communication with backpropagation using buckets. Tune bucket sizing only after profiling; larger buckets can reduce launch overhead but delay overlap.
- Pipeline parallelism: For very large models, split the model into stages on different GPUs and feed it micro-batches, so a stage can compute the next micro-batch while it sends the previous one's activations to the next stage.
For long training runs and multi-node jobs, Lyceum offers reserved capacity: GPUs or a whole cluster held for you and priced per contract, which you reserve with Lyceum's engineers. Ask about the interconnect and location of your cluster before you plan the job.
Cluster Scheduling and Fragmentation
Not every idle GPU is a code problem. On a shared cluster, much of the lost utilization comes from how jobs are scheduled, and it starts with confusing allocation with utilization. Allocation means a GPU is reserved for a user or job; utilization means its streaming multiprocessors are actually executing work. A job can hold eight GPUs while they wait for a network synchronization step or an idle notebook, so the cluster looks full while little compute is happening.
Schedulers also face a bin-packing problem. Most hand out GPUs as whole devices unless sharing is configured; in Kubernetes, for example, a container reserves GPUs through its resource limits. If teams on a 100-GPU cluster request varying amounts, say 3, 5 and 8 GPUs, free capacity quickly fragments. You might have 10 GPUs free, but spread across nodes so that no single node can host an 8-GPU job, and without high-speed interconnects between those nodes they cannot be combined into one large job either.
Static allocation adds to the waste. An engineer who reserves an 8-GPU node for an interactive Jupyter session but runs code only occasionally leaves those GPUs close to 0% utilization for hours. Workload-aware scheduling reduces both problems: frameworks such as Ray let each task or actor declare the GPUs it needs, including a fraction of a GPU for light work, so several small jobs can share one device. Releasing idle interactive reservations and routing light jobs away from the largest GPUs keeps high-demand hardware such as H100s free for compute-bound work.
- Compare allocated with used: Track per-GPU utilization over time, not just how many GPUs are reserved. NVIDIA Data Center GPU Manager (DCGM) collects per-GPU utilization and health metrics that can be aggregated across nodes.
- Find stranded GPUs: Look for free GPUs that no pending job can use because they are scattered across nodes, and pack smaller jobs onto partially used nodes first.
- Audit idle reservations: Set time limits or idle timeouts on interactive sessions so reserved GPUs return to the pool.
Hardware Misalignment: Matching the GPU to the Job
Using the wrong GPU for a task is a reliable way to lower utilization. Fine-tuning a small model on an H100 is often overkill: the model may not have enough parameters, or a large enough batch, to saturate the H100's Tensor Cores, so utilization stays low despite the cost of the instance. The reverse also hurts. Training a large LLM on older GPUs with less memory forces the model across more devices, which adds communication overhead and slows iteration.
The root cause is that teams often do not know a job's resource requirements until they run it, so they over-provision to be safe, and that over-provisioning feeds the utilization gap. A better approach is to estimate memory and compute needs before launch and choose the hardware with the lowest cost per successful training run, not the lowest hourly rate. As a starting point, see how GPU choice differs between a 7B and a 70B model, and use a GPU memory calculator to size VRAM before you book capacity.
Profiling: Stop Guessing, Start Measuring
You cannot fix what you cannot see. If your utilization is low, the first step should always be profiling. NVIDIA Nsight Systems and PyTorch Profiler are the gold standards for this. These tools provide a visual timeline of your execution, showing exactly where the gaps are.
Read the CPU and GPU timelines together. Gaps can indicate data, launch or synchronization stalls; active kernels can still use the device inefficiently. NVIDIA defines nvidia-smi GPU utilization as the percent of time over the past sample period (between 1/6 second and 1 second, depending on the GPU) during which one or more kernels was executing. It is not SM occupancy: it does not show how many streaming multiprocessors were busy or how much math they did, so a single small kernel running back to back can read 100%, and a 100% reading does not imply peak throughput. To see how much of the GPU the work actually uses, read the DCGM profiling metrics SM activity (DCGM_FI_PROF_SM_ACTIVE), SM occupancy (DCGM_FI_PROF_SM_OCCUPANCY) and Tensor Core activity (DCGM_FI_PROF_PIPE_TENSOR_ACTIVE).
For a more granular view, use NVIDIA Nsight Compute. This tool allows you to dive into individual kernels to see if they are memory-bound or compute-bound. If a kernel is memory-bound, it means you are limited by the speed of VRAM; if it is compute-bound, you are limited by the TFLOPS of the cores. This level of detail is essential for researchers optimizing custom CUDA kernels or novel model architectures.
Profiling GPU Utilization on AMD Hardware with ROCm
On AMD Instinct GPUs the same measure-first workflow uses the ROCm profilers. ROCm Systems Profiler (rocprof-sys-run) puts HIP API calls, kernel dispatches, memory copies and GPU telemetry such as utilization on one timeline and correlates them with host call stacks and Python activity, so you can see whether an idle GPU is waiting on a starved data loader or a blocking collective. For lighter tracing, rocprofv3 records kernel dispatches with --kernel-trace, and with --runtime-trace it adds the HIP runtime API and memory operations.
Once you know which kernel matters, ROCm Compute Profiler (rocprof-compute) measures how efficiently a single kernel dispatch uses the hardware: compute unit occupancy, cache and memory bandwidth, and achieved versus peak arithmetic throughput. Older guides call these tools Omniperf and Omnitrace; they were renamed to ROCm Compute Profiler and ROCm Systems Profiler in ROCm 6.3. For a quick live check of utilization, the amd-smi command-line tool monitors GPU performance, much as nvidia-smi does on NVIDIA hardware.
The Lyceum Approach: Prediction and Support
Lyceum Pythia estimates memory and runtime and recommends among T4, A100 and H100 GPUs for supported single-model, single-GPU PyTorch workloads. Its memory prediction leaves out operator overhead from intermediate tensors and data transfer outside the training loop, and its runtime estimate leaves out VM startup and initial data loading, so validate sizing with a representative run. Pythia sizes the hardware; it does not tune distributed training or remove idle time for you.
Lyceum's on-demand GPU VMs run in European data centres and are billed per second with no subscription. Business customers also get a direct line to the engineers who run the platform, with contractual response times.