Throughput
9 articles
Articles
10 June 2026
LLM Tokens Per Second Benchmark: Measured TTFT on Lyceum
This page publishes Lyceum's own measured throughput and first-token latency per model, with the test conditions beside every number, so a workload can be sized against a real measurement rather than a borrowed one.
3 September 2026
Text Embedding API: Throughput, Batching and Cost
Embedding inference is prefill-only, fundamentally changing how workloads scale. Size your corpus backfill and live query path separately, eliminate padding waste, and decide between serverless and dedicated endpoints based on duty cycle rather than instinct.
2 September 2026
Prefill/Decode Disaggregation: Faster Long-Context LLM Serving
Prefill-decode disaggregation splits compute-heavy prompt processing from memory-bound token generation onto separate GPU pools. It eliminates the latency spikes caused when long contexts stall active decodes, optimizing SLO adherence without sacrificing hardware utilization.
1 September 2026
Speculative Decoding: Acceptance Rate vs. Throughput
Speculative decoding trades spare memory bandwidth for faster token generation, but at high concurrency, it competes with real requests and slows down throughput. Here is how to calculate your acceptance rate and find the exact concurrency where your GPU stops being memory-bound.
11 June 2026
vLLM vs TensorRT-LLM: Production Benchmark & Guide
Choosing the right inference engine dictates your infrastructure costs and user experience. We break down the latest performance data to help you optimize your production deployments.
23 April 2026
vLLM Production Deployment Guide: Scaling Sovereign Inference
Moving LLMs from experimental notebooks to production-grade infrastructure requires more than just raw compute. This guide explores how to navigate memory fragmentation, optimize KV caches, and maintain GDPR compliance while scaling vLLM in 2026.
21 April 2026
Reduce LLM Inference Latency on GPUs: A Technical Guide
High latency in LLM inference drives up compute costs and degrades user experience. This guide explores the hardware and software strategies required to minimize Time to First Token (TTFT) and maximize throughput on modern NVIDIA GPUs.
19 April 2026
NVIDIA Dynamo: A Technical Guide to Inference Orchestration
The recent release of NVIDIA Dynamo has fundamentally shifted the landscape for AI infrastructure leads. By bridging the performance gap between open-source frameworks and proprietary engines, this orchestration layer allows teams to maintain full portability without sacrificing throughput.
15 April 2026
Optimizing LLM Inference Throughput with Batching Strategies
Maximizing GPU utilization requires moving beyond simple request-level processing. This guide explores how continuous batching and PagedAttention solve the memory bandwidth bottleneck for production LLM serving.