The Memory-Bandwidth Bottleneck in Inference

Autoregressive generation in large language models is fundamentally memory-bandwidth-bound during the decode phase. To generate a single token, an engine must stream every parameter of the target model from High Bandwidth Memory (HBM) into SRAM and register files. For a 70-billion parameter model in FP16 or BF16 precision, generating one token requires transferring roughly 140 GB of weights across the memory bus. On an NVIDIA H100 SXM with 3.35 TB/s of memory bandwidth, transferring those weights takes approximately 41.8 milliseconds, during which the GPU's 989 TFLOPS of BF16 tensor core compute capacity sits almost entirely idle.

Speculative decoding tackles this arithmetic underutilisation by converting sequential memory reads into parallel verification arithmetic. Instead of running the large target model autoregressively token by token, a smaller draft mechanism proposes a speculative sequence of k candidate tokens. The target model then executes a single forward pass over all k candidate tokens simultaneously. Because evaluating k tokens in parallel reuses the weights already loaded into SRAM, the wall-clock time of this verification pass is nearly identical to generating a single token autoregressively.

This efficiency gain depends strictly on one operational condition: the GPU must have idle compute cycles to spare. Speculative decoding trades arithmetic operations for reduced memory transactions. When those arithmetic units are already saturated, the trade collapses. Understanding whether speculative decoding helps or harms your stack requires looking past optimistic batch-size-one benchmarks and analyzing your hardware utilization profile.

Calculating the Economics of Acceptance Rate

The performance equation of speculative decoding is governed entirely by token acceptance rate. In the original formulation, Leviathan et al. describe speculative decoding as computing several tokens in parallel: an efficient approximation model proposes candidates and the large target model verifies them in a single run, without changing the output distribution. The acceptance rate is the share of those proposed tokens the target model keeps. If a draft model proposes k speculative tokens, the expected number of accepted tokens per verification step, denoted as the acceptance length, follows the geometric sum of accepted prefixes plus one bonus token from the target model's corrected distribution.

Every speculative verification step incurs two distinct computational costs: the sequential forward passes required to generate k draft tokens, and the target model's parallel verification pass over those k tokens. If draft generation takes time t_draft per token and target verification takes time t_target, the total step execution time is (k * t_draft) + t_target. For speculative decoding to deliver a net speedup, the ratio of expected accepted tokens to total step time must exceed the baseline autoregressive generation rate of 1 / t_target.

This arithmetic establishes a strict break-even acceptance rate that you can compute from two numbers you measure yourself. Let r be the draft model's per-token cost expressed as a fraction of one target step (r = t_draft / t_target), and let k be the speculation depth. Generating the draft tokens consumes k * r * t_target, so the total step time is (1 + k * r) * t_target. For the arrangement to deliver a net speedup, the target model must commit an average of more than (1 + k * r) tokens per verification step. If your workload yields an average acceptance length below that threshold, the speculative engine executes more total work while delivering fewer tokens per second than unassisted autoregressive decoding.

  • Low-Entropy Tasks: Code completion, JSON extraction, and structured schema generation exhibit narrow token probability distributions, which is where acceptance rates run highest.
  • High-Entropy Tasks: Creative drafting, open-ended multi-turn reasoning, and complex translation produce wide next-token distributions, causing draft acceptance rates to fall sharply.
  • Draft-to-Target Parameter Sizing: The larger the draft model's latency relative to the target's, the higher the acceptance length you need just to break even, as the formula above makes explicit.
  • KV Cache Memory Overhead: Serving both draft and target models consumes additional VRAM for two distinct sets of KV cache allocations, shrinking the memory headroom required for high-concurrency requests.

Why Vendor Acceptance Benchmarks Tell You Nothing

Hardware and model providers frequently publish speculative decoding benchmarks quoting headline speedup multiples. The original speculative decoding paper, for instance, reports a 2X-3X acceleration on T5-XXL against a standard T5X implementation, with identical outputs. What such headline figures omit is that speculative performance is entirely workload-dependent and measured at a specific batch size on specific hardware. A benchmark measured on synthetic datasets or low-temperature code generation cannot predict how a speculative pipeline will behave on dynamic production traffic.

For example, NVIDIA's Nemotron 3 Super technical report reports an average acceptance length of 3.45 tokens on the SPEED-Bench suite at a fixed draft length of 7, the highest in its comparison set and ahead of DeepSeek-R1's 2.70. Similarly, Z.ai states that GLM-5.2's improved MTP layer increases acceptance length by up to 20%; its published ablation, run on coding scenarios with the number of MTP steps set to 7 for both training and inference, moves acceptance length from a 4.56 baseline to 5.47 once IndexShare and KV Share, rejection sampling and an end-to-end TV loss are stacked. While these figures reflect genuine engineering achievements within specific testing envelopes, they are vendor measurements of their own models rather than universal performance guarantees.

Published acceptance lengths cannot capture the distribution shifts of production environments. User prompts containing domain-specific terminology, foreign language tokens, or high-temperature sampling parameters degrade draft model accuracy. Before committing production infrastructure to a speculative configuration, engineering teams must profile their actual user traffic distributions on raw GPU compute instances to measure empirical acceptance rates.

Speculation StrategyDraft Overhead per StepWhere Acceptance Is HighestNet Latency Impact on Interactive Batches
Independent Small Draft ModelA full forward pass of a second model per draft token, plus its own weights and KV cacheCode and other templated, low-entropy output; weakest on open-ended proseSpeedup at batch size one, provided acceptance clears the break-even threshold
Native Multi-Token Prediction (MTP)One extra head inside the target checkpoint, cheaper than a standalone draft modelStructured and repetitive generation, and whatever the head was trained onSpeedup at batch size one, with no second model to prefill
Prompt Lookup / N-GramCheapest option: an n-gram lookup in the prompt, no extra weights and no extra VRAMSummarisation, document rewriting, RAG quoting and code editing, where output echoes inputStrong speedup on high-overlap prompts, close to neutral elsewhere
Misaligned Draft ModelDraft cost comparable to or above the target step, often with token-mapping overheadNowhere reliably; a mismatched tokenizer or vocabulary depresses acceptance across the boardNet latency loss: the engine does more work per committed token

The High Concurrency Throughput Reversal

The central paradox of speculative decoding is that the exact mechanism that accelerates an isolated request can severely degrade aggregate cluster throughput under heavy load. In production serving engines like vLLM, continuous batching combines incoming requests into dynamic execution batches. At high concurrency, batching naturally saturates the GPU's arithmetic units by aggregating dozens of concurrent memory lookups into dense matrix multiplications.

Once continuous batching shifts the workload from memory-bandwidth-bound to compute-bound, the GPU has no idle tensor core cycles left to offer. Under these conditions, the extra forward passes executed by the draft model and the discarded compute from rejected speculative tokens actively steal execution resources from real user requests. Instead of trading free FLOPs for time, the system burns scarce FLOPs and degrades overall cluster capacity.

This dynamic is explicitly documented by the vLLM team. In official documentation, vLLM scopes speculative decoding strictly to medium-to-low QPS, memory-bound workloads. In a benchmark published on the vLLM engineering blog, running Llama3-70B on 4x NVIDIA H100 GPUs at QPS=1 yielded up to a 1.5x speedup on ShareGPT using a draft model (turboderp/Qwama-0.5B-Instruct) and up to a 2.8x speedup on CNN DailyMail using n-gram prompt lookup. However, at high QPS on the same hardware, the same blog records a 1.4x slowdown on ShareGPT and a 1.8x slowdown on CNN DailyMail, because the extra compute required to propose and verify tokens degrades an already compute-bound system.

Finding Where Your Server Stops Being Memory-Bound

Because the crossover point where speculative decoding becomes a net penalty depends entirely on your specific model, hardware, sequence lengths, and traffic concurrency, you cannot rely on external heuristics. You must locate your deployment's inflection point through systematic empirical profiling.

To identify this threshold, execute a controlled concurrency sweep using your production traffic replay. Measure two distinct metrics simultaneously: per-request Time Per Output Token (TPOT) and aggregate server-wide throughput (tokens per second per GPU). Sweep your request concurrency from batch size 1 up to maximum queue saturation, testing both baseline autoregressive decoding and speculative decoding across identical request distributions.

  1. Step 1: Deploy a baseline instance with standard batching strategies and speculative decoding disabled. Record TPOT and aggregate output tokens/sec across increasing concurrency tiers (1, 2, 4, 8, 16, 32, 64 concurrent streams).
  2. Step 2: Deploy an identical instance with your chosen speculative decoding method enabled (draft model, MTP head, or prompt lookup). Replay the identical prompt distribution across the same concurrency tiers.
  3. Step 3: Plot the two curves. At low concurrency, the speculative configuration will demonstrate lower TPOT. As concurrency increases, the TPOT curves will converge while aggregate throughput curves will cross.
  4. Step 4: Identify the crossing concurrency. If your median operational concurrency sits below this crossing point, speculative decoding delivers net value. If your production cluster regularly operates above it, disable speculation to maximize global throughput.

Running these benchmarks requires dedicated compute environments where memory bandwidth and GPU clock frequencies remain completely isolated from noisy neighbours. Provisioning an On-demand GPU VM provides raw SSH access to bare-metal NVIDIA H100 or A100 instances with NVLink interconnects, allowing engineering teams to run precise parameter sweeps and capture accurate telemetry before deploying serving configurations to production.

The Prefill Tax: Speculation Does Not Accelerate TTFT

A frequent misunderstanding in production deployment is assuming that speculative decoding improves end-to-end API response times across all query types. Speculative decoding operates exclusively on the autoregressive decode phase. It does nothing to accelerate the initial prefill phase, meaning it provides zero reduction in Time to First Token (TTFT).

Furthermore, speculative decoding introduces a direct prefill penalty. When using an independent draft model, the draft architecture must also process the entire input prompt context to initialize its own KV cache before it can propose candidate tokens. For long prompts, running dual prefill passes across both the target model and the draft model introduces measurable initial latency overhead, delaying the generation of the first token.

For workloads characterised by large prompt contexts and short generation targets, such as document extraction, enterprise RAG, and classification, this prefill tax easily outweighs any minor decode acceleration. When a pipeline processes a long prompt to generate only a short summary, most of the request's execution time sits in the prefill stage, which speculation does not touch at all. Adding draft prefill latency in front of a decode phase that lasts only a handful of tokens produces a net slowdown across the total request lifecycle.

  • Evaluate Input-to-Output Ratio: If median prompt length exceeds output length by more than 10:1, speculative decoding is rarely economical.
  • Monitor TTFT Degradation: Track p95 TTFT metrics; dual-model prefill directly inflates latency for time-sensitive interactive applications.
  • Consider Architectural Disaggregation: For workloads with severe prefill bottlenecks, decoupling prefill and decode execution across specialized hardware nodes yields far greater latency reductions than speculative token drafting alone.

Draft-Free Methods and Open Inference Stacks

To bypass the computational overhead and operational fragility of hosting separate draft models, recent advancements in inference engineering focus on draft-free and integrated head speculation techniques. A primary limitation of standalone draft models is the strict requirement for identical tokenizers. If a small model uses a different vocabulary or token merging strategy than the target model, integrating them requires complex token mapping layers that degrade acceptance rates.

Prompt lookup decoding eliminates auxiliary models entirely by extracting candidate n-grams directly from the input prompt. Because language generation frequently echoes terms and structures present in the context, prompt lookup achieves high acceptance rates on summarisation, document rewriting, and code editing without consuming additional VRAM or requiring auxiliary weights. For model-level speculation, EAGLE-3 replaces EAGLE's reliance on top-layer features with multi-layer feature fusion via a technique its authors call training-time test, moving to direct token prediction so the drafter keeps benefiting from more training data; the paper reports speedup ratios up to 6.5x and a 1.38x throughput improvement at batch size 64 in SGLang. Similarly, next-generation open weights such as DeepSeek-V4-Pro integrate native Multi-Token Prediction (MTP) modules directly into the primary checkpoint, avoiding auxiliary memory allocation. Older methods like Medusa remain functional configuration flags in the vLLM engine, though production deployment has largely transitioned to MTP and EAGLE implementations.

Navigating these low-level serving optimizations requires complete architectural transparency. Proprietary inference platforms like Together and Fireworks operate black-box serving engines that conceal internal batching heuristics, speculation depths, and scheduling parameters behind opaque API abstractions. When deploying mission-critical infrastructure, AI teams need full visibility into the underlying serving stack to optimize token economics and maintain control over hardware execution.

At Lyceum, we build European AI infrastructure on an open, transparent inference stack powered by vLLM, NVIDIA Dynamo, and TensorRT-LLM. Whether you consume pre-hosted open models via Serverless Inference with OpenAI SDK compatibility, deploy private pipelines on Dedicated Inference, or profile custom draft architectures on raw On-demand GPU VM instances in our European data centres, you maintain uncompromised control over your serving parameters, data residency, and computational efficiency.