Inference Serving

Running models in production on serverless infrastructure: deployment, scaling, cold starts, batch, rate limits and custom model hosting. Serves teams moving from evaluation to production.

36 articles

Articles

7 October 2026

Cutting Inference Spend by Routing Requests by Difficulty

Some requests may work well on a smaller model. Measure that share, include verification and escalation costs, and test quality before routing production traffic.

23 February 2026

KV Cache Memory Calculation for LLMs: A Technical Guide

Large Language Model (LLM) weights are only half the story. As sequence lengths grow and batch sizes increase, the Key-Value (KV) cache often becomes the primary consumer of GPU VRAM, leading to the dreaded Out-of-Memory (OOM) errors that plague production environments. For ML engineers, understanding the precise memory requirements of the KV cache is not just a theoretical exercise; it is a prerequisite for efficient scaling. This article provides a deep dive into the mechanics of KV caching, the mathematical foundations for memory estimation, and how modern architectures like Llama 3 or Mistral utilize advanced attention mechanisms to mitigate memory bottlenecks.

10 June 2026

LLM Tokens Per Second Benchmark: Measured TTFT on Lyceum

This page publishes Lyceum's own measured throughput and first-token latency per model, with the test conditions beside every number, so a workload can be sized against a real measurement rather than a borrowed one.

14 August 2026

How to Use Lyceum API within Claude Code

This guide shows how to run Claude Code on an open-weight model through Serverless Inference, with three commands and no proxy, and how to check which model is answering.

17 September 2026

Open Inference Stack: vLLM, Dynamo vs Proprietary Engines

If you are weighing an open serving stack against a proprietary engine, this guide separates the speed question from the portability question. We examine what vLLM and NVIDIA Dynamo buy you, keeping open source software, self-hosted versus managed deployments, and exposed controls as separate axes to evaluate.

31 August 2026

Autoscaling GPU Inference: KServe vs Ray Serve vs llm-d

KServe, Ray Serve, and llm-d offer different approaches to scaling GPU inference on Kubernetes. While KServe standardizes general model serving and Ray Serve enables Python-native pipelines, llm-d adds LLM-specific optimizations like disaggregated prefill and decode.

18 September 2026

Time to First Token: What Actually Determines LLM API Latency

Time to first token dictates how fast your AI product feels to users. This technical breakdown explores the infrastructure layers that drive LLM latency, from queueing and continuous batching to prompt prefilling and Server-Sent Events.

10 September 2026

vLLM vs SGLang vs TensorRT-LLM (2026): Picking a Serving Engine

Comparing vLLM, SGLang, and TensorRT-LLM on peak throughput is the wrong approach. The real variables that dictate inference performance are model churn and prefix sharing - and for most platform teams, the most practical solution is to decline the engine choice entirely.

9 September 2026

Static IP and DNS for GPU Inference Endpoints

Enterprise buyers often request a static IP and custom DNS for GPU inference endpoints to satisfy default-deny egress firewalls. This guide breaks down why managed endpoints rarely offer static IPs, how to structure egress policies securely, and the connectivity questions to ask providers. For Dedicated Inference and large workloads, contact a Lyceum engineer to review custom dedicated deployment options.

3 September 2026

Text Embedding API: Throughput, Batching and Cost

Embedding inference is prefill-only, fundamentally changing how workloads scale. Size your corpus backfill and live query path separately, eliminate padding waste, and decide between serverless and dedicated endpoints based on duty cycle rather than instinct.

2 September 2026

Prefill/Decode Disaggregation: Faster Long-Context LLM Serving

Prefill-decode disaggregation splits compute-heavy prompt processing from memory-bound token generation onto separate GPU pools. It eliminates the latency spikes caused when long contexts stall active decodes, optimizing SLO adherence without sacrificing hardware utilization.

1 September 2026

Speculative Decoding: Acceptance Rate vs. Throughput

Speculative decoding trades spare memory bandwidth for faster token generation, but at high concurrency, it competes with real requests and slows down throughput. Here is how to calculate your acceptance rate and find the exact concurrency where your GPU stops being memory-bound.

31 August 2026

Fixing vLLM CUDA Out of Memory: KV Cache Tuning Guide

vLLM CUDA out of memory errors usually stem from startup reservations, not runtime loads. By tuning gpu_memory_utilization and max-model-len, you can right-size the KV cache and stabilize inference without renting larger GPUs.

15 April 2026

Serverless GPU Inference: Architecture, Economics, and Compliance

11 June 2026

vLLM vs TensorRT-LLM: Production Benchmark & Guide

Choosing the right inference engine dictates your infrastructure costs and user experience. We break down the latest performance data to help you optimize your production deployments.

10 June 2026

Serverless GPU Cold Start Latency: Architecture Comparison

Scale-to-zero GPU infrastructure promises massive cost savings, but a 40-second cold start will kill any real-time AI application. Here is a technical breakdown of where the time actually goes and how modern inference stacks are solving the VRAM bottleneck.

4 June 2026

Long Context Inference: GPU Requirements & VRAM Guide

Context kills VRAM. Learn the exact math behind KV cache bottlenecks and how to architect your GPU infrastructure for 128K+ token workloads.

31 May 2026

LLM Context Length vs. GPU Memory: Calculating VRAM Requirements

Parameter count only tells half the story. Learn how to calculate the exact GPU memory required for long-context LLM inference and avoid catastrophic Out-of-Memory errors in production.

30 May 2026

The Guide to Serving Fine-Tuned LLMs in Production

Training a model is no longer the hard part. Serving fine-tuned models at scale requires avoiding memory bottlenecks and excessive costs for idle GPUs.

28 May 2026

Deploy a Hugging Face Model Inference API: 2026 Production Guide

Moving a Hugging Face model from a local notebook to a production API requires solving three hard problems: GPU memory fragmentation, unpredictable cold starts, and strict data residency requirements.

24 May 2026

Deploy Hugging Face Model to GPU Cloud

Moving a Hugging Face model from a local notebook to production requires strict VRAM math and the right inference engine. Learn how to deploy open-source LLMs at scale without hyperscaler cost overruns.

18 May 2026

GGUF vs GPTQ vs AWQ: The Definitive LLM Quantization Framework

We break down the exact performance, memory, and throughput differences between GGUF, GPTQ, and AWQ for production inference.

23 April 2026

Serverless Inference Cold Start Latency: A Technical Optimization Guide

Cold starts remain the primary barrier to responsive serverless AI. This guide breaks down the technical stages of GPU initialization and provides a framework for minimizing latency in production environments.

23 April 2026

vLLM Production Deployment Guide: Scaling Sovereign Inference

Moving LLMs from experimental notebooks to production-grade infrastructure requires more than just raw compute. This guide explores how to navigate memory fragmentation, optimize KV caches, and maintain GDPR compliance while scaling vLLM in 2026.

22 April 2026

Self-Host LLM APIs on EU Infrastructure: The Modern Guide

As hyperscaler credits expire and the EU AI Act's high-risk obligations phase in, deferred to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I systems, AI teams are moving toward sovereign infrastructure. This guide explores how to self-host LLM APIs in Europe to ensure data residency without sacrificing performance.

21 April 2026

Reduce LLM Inference Latency on GPUs: A Technical Guide

High latency in LLM inference drives up compute costs and degrades user experience. This guide explores the hardware and software strategies required to minimize Time to First Token (TTFT) and maximize throughput on modern NVIDIA GPUs.

21 April 2026

The Economics of Scale to Zero: Slashing GPU Inference Costs in 2026

Running dedicated GPU instances for bursty inference workloads is the fastest way to burn through venture capital. Scale-to-zero orchestration allows teams to eliminate idle compute costs without sacrificing the performance required for production-grade AI.

19 April 2026

Multi-Model Serving on Single GPUs with vLLM and PagedAttention

Dedicating a high-end GPU to a single model often leaves most of the card idle and the unit economics unsustainable. Modern inference stacks now allow for concurrent model execution on a single H100 or B200 node without the latency penalties of traditional context switching.

19 April 2026

NVIDIA Dynamo: A Technical Guide to Inference Orchestration

The recent release of NVIDIA Dynamo has fundamentally shifted the landscape for AI infrastructure leads. By bridging the performance gap between open-source frameworks and proprietary engines, this orchestration layer allows teams to maintain full portability without sacrificing throughput.

18 April 2026

Host Fine-Tuned Model Production APIs: A Technical Guide

Moving a fine-tuned model from a local notebook to a production API requires solving for memory management, cold starts, and unsustainable hyperscaler costs. This guide explores the technical architecture needed to serve LLMs with high throughput while keeping processing inside European data centers.

18 April 2026

Self-Hosted LLM API Gateway Guide: Architecture and Infrastructure

Fragmented model access often leads to security vulnerabilities and unpredictable cost overruns. A self-hosted LLM API gateway centralizes control, ensuring GDPR compliance while providing a unified interface for your inference workloads.

17 April 2026

Deploying Mistral Large on European GPU Cloud Infrastructure

European AI teams face a dilemma: high-performance LLMs like Mistral Large 2 require massive GPU clusters, but US-based clouds often fail strict GDPR and data residency requirements. This guide explores how to deploy Mistral Large 2 on EU-sovereign infrastructure without the hyperscaler price tag.

17 April 2026

Deploying Private LLM Endpoints on GPU Cloud: A 2026 Strategy

As AI startups outgrow their initial cloud credits, the shift toward private LLM endpoints becomes a necessity for cost control and GDPR compliance. This guide examines the technical architecture and economic frameworks required to deploy high-performance inference on European GPU infrastructure.

16 April 2026

Deploying Custom Docker Model Inference APIs for Production

Moving beyond black-box APIs requires a robust containerization strategy and optimized GPU orchestration. This guide explores how to build and deploy custom Docker inference endpoints that maintain data residency while maximizing throughput.

16 April 2026

Deploying Llama 3 Inference APIs on Sovereign GPU Clouds

Scaling Llama 3 inference requires balancing VRAM bottlenecks against unsustainable hyperscaler costs. This guide explores how to deploy production-grade APIs using European infrastructure and modern orchestration stacks.

15 April 2026

Optimizing LLM Inference Throughput with Batching Strategies

Maximizing GPU utilization requires moving beyond simple request-level processing. This guide explores how continuous batching and PagedAttention solve the memory bandwidth bottleneck for production LLM serving.

Your next workload starts here