Deconstructing RAG Pipeline Query Costs

Why does a Retrieval-Augmented Generation (RAG) pipeline that costs fractions of a cent in development quickly escalate into tens of thousands of dollars in monthly cloud invoices once deployed? For engineering teams scaling production search and knowledge retrieval, cost predictability is broken by treating a multi-stage pipeline as a single API call. In practice, calculating the true unit economics of a RAG query requires profiling three distinct compute stages independently: embedding generation, vector retrieval, and Large Language Model (LLM) answer generation.

The Three Compute Domains of RAG

Each stage in a RAG query stresses distinct hardware resources across compute, memory bandwidth, and memory capacity. The initial stage converts the user query into dense vector representations. This phase is compute-bound but uses minimal VRAM, executing rapidly on GPU Tensor Cores. The second stage performs Approximate Nearest Neighbor (ANN) search across a vector index to retrieve relevant context chunks, which is primarily memory-bound or index-bound. The final and most expensive stage is LLM generation, where retrieved chunks are concatenated with the prompt and passed to the model for autoregressive decoding.

Pipeline StagePrimary Hardware ResourceTypical VRAM FootprintLatency ProfileShare of Query Compute
Embedding GenerationGPU Tensor Cores (FP16/BF16)1 GB to 4 GB5 ms to 20 ms< 2%
Vector RetrievalHost RAM / GPU VRAM (cuVS/CAGRA)Variable (Index Size)2 ms to 15 ms< 3%
LLM Prompt PrefillGPU Tensor Cores (Compute-Bound)Model Weights + Input KV30 ms to 150 ms15% to 25%
LLM Token DecodingGPU Memory Bandwidth (HBM)Model Weights + Full KV Cache200 ms to 1200 ms70% to 80%

When engineering teams look at their end-to-end latency and compute spend, LLM decoding accounts for over three-quarters of the entire cost. Sizing RAG pipeline infrastructure accurately requires isolating these phases rather than applying blanket per-token estimates.

Embedding Generation and Retrieval Benchmarks

The embedding phase is often over-engineered on dedicated infrastructure when lightweight serving is all that is required. An embedding model like an 8-billion parameter encoder or modern dense embedding architecture typically requires between 1 GB and 4 GB of VRAM when loaded in half-precision (FP16 or BF16). Because embedding generation processes short input queries in a single forward pass, it achieves high throughput without accumulating sequential state.

Evaluating Embedding Models on MTEB

Engineering teams select embedding backbones using the Massive Text Embedding Benchmark (MTEB), which spans 8 embedding tasks covering 58 datasets and 112 languages, including retrieval, clustering, and reranking. Models near the top of the public MTEB leaderboard deliver strong NDCG@10 retrieval accuracy while maintaining modest compute footprints.

For high-volume production retrieval, a serverless embedding endpoint such as Qwen3 Embedding 8B is listed at $0.01 per 1M input tokens, so embedding short user queries costs a rounding error relative to generation. Even when factoring in periodic document re-indexing, the embedding step remains a minor fraction of the total RAG cost equation.

  • Query embedding generation: short queries are billed at the published input-token rate of $0.01 per 1M tokens, a negligible per-query amount.
  • VRAM overhead: Small footprint (1 GB to 4 GB) allows colocation on smaller GPUs or shared instances.
  • Throughput characteristics: Pure matrix multiplication during prefill scales linearly with batch size without sequential decoding penalties.

Where the Bottleneck Sits: Prefill vs. Decode

The core economic and computational bottleneck in any RAG pipeline sits squarely in the LLM generation phase. When context is retrieved from the vector database, it inflates the Input Sequence Length (ISL). A standard retrieval-augmented generation pipeline passes roughly 3,000 input tokens of context to generate a 500-token response. In practice that budget is made up of the user prompt, the top retrieved chunks, and a system prompt carrying instructions and formatting rules.

Compute-Bound Prefill vs. Memory-Bandwidth-Bound Decode

LLM execution splits into two radically different mathematical phases: prefill and decode. In the prefill phase, the inference engine processes all 3,000 input tokens concurrently. This phase is compute-bound, saturating GPU Tensor Cores at high mathematical intensity (TFLOPS) to compute key and value vectors. The prefill latency (Time-to-First-Token, or TTFT) is relatively low per token processed.

In contrast, the decode phase generates output tokens sequentially, one token at a time. To generate a single output token, the GPU must load every parameter of the model weights from High Bandwidth Memory (HBM) into on-chip cache and registers. If you generate a 500-token answer with an 8-bit quantized 70B parameter model, the GPU memory bus must read approximately 70 GB of weights 500 consecutive times. The decode phase is strictly memory-bandwidth bound, causing token generation latency (Time-Per-Output-Token, or TPOT) to accumulate and dominating the final query cost.

The KV Cache and API Generation Markups

As the LLM decodes each new token, it must retain the Key and Value representations of all previous tokens (both input prompt and previously generated output tokens) in VRAM. This Key-Value (KV) cache grows dynamically with sequence length and concurrency. In traditional inference setups without virtual memory management, systems preallocate large contiguous memory blocks based on maximum context lengths, and existing systems waste 60% to 80% of that memory through fragmentation and over-reservation.

Why Cloud APIs Charge an Output Token Premium

Public cloud API providers face this severe memory-bandwidth bottleneck and pass the cost of idle tensor cores and reserved KV cache directly to customers. Across major proprietary and hyperscaler API platforms, output tokens commonly cost four to eight times more than input tokens. For example, Google Vertex AI lists Gemini 2.5 Pro at USD 1.25 per million input tokens against USD 10.00 per million output tokens (an 8:1 ratio), and Amazon Bedrock lists Claude 3.5 Sonnet on extended access at USD 6.00 input against USD 30.00 output (a 5:1 ratio).

On open-weight models, hyperscaler pricing narrows the gap slightly while maintaining a premium: Amazon Bedrock lists Llama 2 Chat 70B on demand at USD 1.95 per 1M input tokens against USD 2.56 per 1M output tokens. In a RAG pipeline producing long answers at sustained daily volume, this retail markup multiplies generation costs.

  • PagedAttention memory management: Dynamically allocates KV cache in non-contiguous physical memory pages, cutting memory waste to under 4% against the 60% to 80% wasted by contiguous preallocation.
  • Eliminating fragmentation: Reclaiming wasted VRAM lets the system batch more sequences together on the same physical hardware, raising GPU utilization and throughput.
  • Direct cost capture: Running on dedicated GPUs allows teams to bypass retail API output token premiums and pay only for actual hardware execution time.

LLM Batching and the Underutilization Penalty

A frequent pitfall when sizing RAG infrastructure is assuming full steady-state GPU utilization. When traffic is sparse or requests arrive sequentially, a dedicated GPU spends the majority of its time idling between memory transfers. If your pipeline receives a single request per second on an instance capable of serving many times that load, a cost calculation based on peak throughput badly understates your effective cost per query. There is no universal break-even utilization: work it out from the hourly rate you are quoted, the tokens per second you sustain, and your current per-token rate.

Continuous Batching and Concurrency Scaling

To maximize efficiency, production inference engines implement continuous batching (iteration-level scheduling). Instead of waiting for an entire batch to finish generating before accepting new queries, continuous batching injects new prefill requests into the execution cycle at each decoding step. This amortizes the cost of streaming model weights from HBM across dozens of concurrent requests.

The effect of continuous batching is easiest to reason about as a curve rather than a table. At very low offered load (roughly one request per second) the scheduler only ever holds one or two sequences in flight, so each decode step streams the full model weights from HBM to serve almost no work, and the cost per query is at its worst. As load rises and the resident batch grows to a handful and then to a few dozen concurrent sequences, that same weight-streaming cost is shared across every stream in the batch and the cost per query falls steeply. Near saturation the GPU is running the largest batch its KV cache can hold, and cost per query bottoms out. Because the exact utilization and crossover point depend on your model size, context length, and hourly rate, measure them on your own traffic rather than borrowing a vendor's saturated-throughput number.

Managing GPU idle time and matching instance capacity to traffic patterns is essential. When traffic drops below the break-even point of dedicated infrastructure, combining serverless inference for baseline bursts with dedicated nodes for base load stabilizes unit economics.

Sizing the GPU: H100 vs. B200 Economics

Selecting the right accelerator for RAG generation requires evaluating cost per million tokens and memory bandwidth rather than focusing solely on hourly rental rates. Sizing choices directly dictate how many concurrent 3,000-token context windows fit into VRAM before running out of memory.

Comparing NVIDIA H100 and B200 for Production Workloads

The NVIDIA H100 has been the standard production baseline, offering 80 GB of HBM3 memory and 3.35 TB/s of memory bandwidth. For a 70-billion parameter model quantized to FP8 (requiring ~70 GB for weights), an 8-GPU H100 node provides 640 GB of aggregate VRAM. This leaves approximately 570 GB for KV cache allocation, supporting substantial concurrent batch sizes across standard RAG context windows.

The NVIDIA B200 introduces 192 GB of HBM3e memory per GPU with memory bandwidth reaching up to 8.0 TB/s. With 192 GB VRAM per card, a single dual-GPU setup or an 8-GPU node provides 1.5 TB of aggregate high-speed memory. The increased memory bandwidth accelerates the sequential decode phase, while the massive memory capacity allows serving four times as many concurrent active KV contexts without offloading.

  • NVIDIA H100 (80 GB VRAM, 3.35 TB/s bandwidth): High compute throughput, optimal for mid-tier concurrency and standard context lengths.
  • NVIDIA H200 (141 GB VRAM, 4.8 TB/s bandwidth): Expanded HBM3e capacity, enabling longer retrieved context chunks without tensor parallelism overhead.
  • NVIDIA B200 (192 GB VRAM, up to 8.0 TB/s bandwidth): High-density architecture designed to maximize batch concurrency and reduce per-token serving costs at enterprise scale.

Scaling to Sovereign Dedicated Infrastructure

When transitioning from prototyping to high-throughput production, moving from pay-per-token APIs to dedicated infrastructure gives engineering teams complete control over their latency, batching pipelines, and unit economics. Choosing whether to deploy right-size GPU instances via On-demand GPU VM or provision a Large-Scale GPU Cluster depends directly on your daily sustained query volume.

European Data Sovereignty and Total Cost Optimization

For engineering teams operating in Europe, infrastructure architecture involves compliance and data governance alongside hardware performance. Lyceum operates dedicated GPU infrastructure located in European data centres (including France and the Nordic regions), providing zero data retention, strict GDPR compliance, and complete architectural isolation from the US CLOUD Act. This ensures that proprietary enterprise documents retrieved during RAG execution never leave EU jurisdiction.

By sizing each stage of your RAG pipeline independently (offloading embeddings to lightweight endpoints, utilizing continuous batching with PagedAttention for generation, and right-sizing dedicated NVIDIA H100 or B200 clusters), you eliminate retail token markups and establish predictable cost per query at scale.