Endpoint Types
13 articles
Articles
14 August 2026
How to Use Lyceum API within Claude Code
This guide shows how to run Claude Code on an open-weight model through Serverless Inference, with three commands and no proxy, and how to check which model is answering.
31 August 2026
Autoscaling GPU Inference: KServe vs Ray Serve vs llm-d
KServe, Ray Serve, and llm-d offer different approaches to scaling GPU inference on Kubernetes. While KServe standardizes general model serving and Ray Serve enables Python-native pipelines, llm-d adds LLM-specific optimizations like disaggregated prefill and decode.
10 September 2026
vLLM vs SGLang vs TensorRT-LLM (2026): Picking a Serving Engine
Comparing vLLM, SGLang, and TensorRT-LLM on peak throughput is the wrong approach. The real variables that dictate inference performance are model churn and prefix sharing - and for most platform teams, the most practical solution is to decline the engine choice entirely.
9 September 2026
Static IP and DNS for GPU Inference Endpoints
Enterprise buyers often request a static IP and custom DNS for GPU inference endpoints to satisfy default-deny egress firewalls. This guide breaks down why managed endpoints rarely offer static IPs, how to structure egress policies securely, and the connectivity questions to ask providers. For Dedicated Inference and large workloads, contact a Lyceum engineer to review custom dedicated deployment options.
15 April 2026
Serverless GPU Inference: Architecture, Economics, and Compliance
28 May 2026
Deploy a Hugging Face Model Inference API: 2026 Production Guide
Moving a Hugging Face model from a local notebook to a production API requires solving three hard problems: GPU memory fragmentation, unpredictable cold starts, and strict data residency requirements.
24 May 2026
Deploy Hugging Face Model to GPU Cloud
Moving a Hugging Face model from a local notebook to production requires strict VRAM math and the right inference engine. Learn how to deploy open-source LLMs at scale without hyperscaler cost overruns.
22 April 2026
Self-Host LLM APIs on EU Infrastructure: The Modern Guide
As hyperscaler credits expire and the EU AI Act's high-risk obligations phase in, deferred to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I systems, AI teams are moving toward sovereign infrastructure. This guide explores how to self-host LLM APIs in Europe to ensure data residency without sacrificing performance.
18 April 2026
Host Fine-Tuned Model Production APIs: A Technical Guide
Moving a fine-tuned model from a local notebook to a production API requires solving for memory management, cold starts, and unsustainable hyperscaler costs. This guide explores the technical architecture needed to serve LLMs with high throughput while keeping processing inside European data centers.
17 April 2026
Deploying Mistral Large on European GPU Cloud Infrastructure
European AI teams face a dilemma: high-performance LLMs like Mistral Large 2 require massive GPU clusters, but US-based clouds often fail strict GDPR and data residency requirements. This guide explores how to deploy Mistral Large 2 on EU-sovereign infrastructure without the hyperscaler price tag.
17 April 2026
Deploying Private LLM Endpoints on GPU Cloud: A 2026 Strategy
As AI startups outgrow their initial cloud credits, the shift toward private LLM endpoints becomes a necessity for cost control and GDPR compliance. This guide examines the technical architecture and economic frameworks required to deploy high-performance inference on European GPU infrastructure.
16 April 2026
Deploying Custom Docker Model Inference APIs for Production
Moving beyond black-box APIs requires a robust containerization strategy and optimized GPU orchestration. This guide explores how to build and deploy custom Docker inference endpoints that maintain data residency while maximizing throughput.
16 April 2026
Deploying Llama 3 Inference APIs on Sovereign GPU Clouds
Scaling Llama 3 inference requires balancing VRAM bottlenecks against unsustainable hyperscaler costs. This guide explores how to deploy production-grade APIs using European infrastructure and modern orchestration stacks.