Model Library
Reference coverage of every model in the Inference Studio catalogue: capabilities, context windows, hosting region and per-token prices. Serves buyers evaluating a specific open-weight model.
39 articles
Articles
15 June 2026
DeepSeek-V4-Pro: specs, benchmarks, and how to run it on Lyceum
DeepSeek-V4-Pro delivers frontier-level reasoning and a massive 1M-token context window. Learn how to deploy it through Lyceum's OpenAI-compatible API with simple per-token pricing.
31 July 2026
Kimi K3 API: Where to Run It, and What a 1M-Token Context Costs
Kimi K3 is Moonshot AI's open-weights model with a 1M-token context window. On Lyceum it runs as moonshotai/kimi-k3 at $3.00 input, $0.75 cached input and $15.00 output per million tokens, EU-hosted, with prompts and outputs not stored or used for training. Kimi K2.6 was retired on 5 October 2026, and its model id now routes to Kimi K3.
8 September 2026
GLM-5.2 Instant: Latency-Optimised GLM, EU-Hosted
GLM-5.2 Instant offers the same 1M-token context window and identical per-token pricing as the base GLM-5.2 model, but is positioned for lower latency. It runs on EU-sovereign serverless inference in eu-north1, so teams can evaluate throughput on their own traffic.
11 September 2026
EU Hosted LLM API: Lyceum Model Roster, Prices and Changes
Which model strings answer on this OpenAI-compatible endpoint, which offering records publish a hosting location and complete price data, and which older strings are absent from the live roster. Facts checked on 10 September 2026, with a roster request you can rerun yourself.
1 September 2026
GLM-5.3: specs, benchmarks, and how to run it on Lyceum
GLM-5.3 is ZAI's latest MoE model, offering a 1M-token context window and emergent cyber capabilities for agentic workflows. It is available on Serverless Inference via a drop-in OpenAI-compatible API, billed purely per token.
30 September 2026
Where to Run DeepSeek in Europe: V4 Flash vs Pro
If DeepSeek is blocked at your compliance review, this page answers per model where V4 Flash and V4 Pro run in Europe and what each costs.
18 September 2026
DeepSeek-V4.1-Flash: specs, benchmarks, and how to run it
DeepSeek-V4.1-Flash is live on Lyceum Serverless Inference: this page covers what changed, what it costs per token ($0.50 input, $0.13 cached, $1.50 output per 1M) and the exact request to send. Architecture and benchmark figures are DeepSeek's own, vendor-reported.
8 September 2026
Qwen3.5-9B: Compact Qwen With the Narrowest Price Spread
Qwen3.5-9B is a compact open-weight model with a 256K context window and the narrowest input-to-output price spread in the catalogue. Costing $0.15 per million input tokens and $0.20 per million output tokens, it significantly reduces the total bill for output-heavy workloads.
25 August 2026
Running GLM 5.1, 5.2 and 5.2 Instant in Europe: Self-Hosting and Serverless Options
Z.ai's GLM-5 series introduces 1M-token contexts and powerful agentic capabilities via a 744B MoE architecture. For European teams, running these models locally requires massive GPU clusters, making a managed serverless endpoint a highly practical alternative.
12 August 2026
Where to Run Kimi Models in Europe: K2.6, K2.7 Code and K3
Moonshot AI's Kimi models deliver frontier capabilities for agentic coding. K2.6 and K2.7 Code offer 1T-parameter scale with 256K context, while K3 pushes to 2.8T parameters and a 1M-token window. European teams can run them via EU-hosted APIs to maintain data residency.
27 June 2026
GLM-5.2: specs, benchmarks, and how to run it on Lyceum
7 September 2026
MiniMax-M3: EU-Hosted Where M2.5 Is Global
If you are evaluating MiniMax under a data-residency requirement, this page states which of the two models has a confirmed European region and what choosing it costs you per token.
7 September 2026
Kimi-K2.7-Code: Code-Specialised Kimi, EU-Hosted
Kimi-K2.7-Code brings a 256K context and a confirmed EU region to open-weight coding models. With input priced at $1.25 and output at $4.50 per million tokens, understanding this ratio is critical for managing the cost of agentic software engineering.
1 September 2026
Qwen3.8 Flash Next: specs, benchmarks and Lyceum API
Qwen3.8 Flash Next is a multimodal MoE model previewing the Qwen4 architecture, activating just 6B parameters per token for high-efficiency agent workflows. It is available on Lyceum Serverless Inference via an OpenAI-compatible API with no infrastructure overhead.
1 September 2026
Qwen3.8 27B: specs, benchmarks, and how to run it on Lyceum
Qwen3.8 27B is a 27-billion-parameter dense multimodal model offering a 256K context window. Now available on Lyceum Serverless Inference, it supports prompt caching and built-in reasoning traces for complex agentic workloads at $0.40 per 1M input tokens.
1 September 2026
Qwen3.8 2.4T A95B: specs, benchmarks, and how to run it on Lyceum
Qwen3.8 2.4T A95B is the new open-weight flagship, featuring a 256K context window and a hybrid-attention MoE architecture. It is available now on Lyceum Serverless Inference via an OpenAI-compatible API, billed per token with zero provisioning overhead.
1 September 2026
GLM-5.3 Flash: specs, benchmarks, and how to run it on Lyceum
GLM-5.3 Flash is ZAI's highly efficient MoE model, featuring 18B active parameters and a 1M-token context window. Available now on Lyceum Serverless Inference, it delivers frontier coding and agentic capabilities starting at $0.05 per 1M cached input tokens.
20 August 2026
DeepSeek-V4-Flash: specs, benchmarks, and how to run it
DeepSeek-V4-Flash is a 284-billion parameter MoE model offering agentic reasoning across a 1-million token context window. Lyceum serves it via an OpenAI-compatible API from eu-north1 in the European Union, optimized for enterprise inference at $0.15 per million input tokens.
1 August 2026
DeepSeek V4 Flash: 1M-Token Context for AI Products
DeepSeek V4 Flash introduces a 284B parameter MoE architecture with 13B active parameters, delivering low time-to-first-token latency and a 1,048,576-token context window. For AI-native products, this means high-throughput agent loops and long-context retrieval hosted natively in Europe
14 August 2026
DeepSeek V4 Pro API: EU Hosting, Pricing and Context Limits
DeepSeek V4 Pro API runs in European data centres with 1M token context, $2.00/$4.00 pricing per 1M tokens, zero data retention, and full OpenAI SDK compatibility.
16 May 2026
Deploying Microsoft Phi-4 Inference on GPU Cloud: A Production Guide
27 June 2026
Qwen3-Embedding-8B: specs, benchmarks, and how to run it on Lyceum
Qwen3-Embedding-8B delivers state-of-the-art retrieval performance across 100+ languages. Built on the Qwen3 foundation, it supports customizable output dimensions and instruction-aware queries for complex RAG pipelines.
27 June 2026
Wan Image: specs, benchmarks, and how to run it on Lyceum
Wan Image delivers photorealistic generation with advanced prompt adherence. Here is how to deploy it on Lyceum Technology.
25 June 2026
Qwen3-235B-A22B: specs, benchmarks, and how to run it on Lyceum
Qwen3-235B-A22B-Instruct-2507 is Alibaba's flagship Mixture-of-Experts model, activating only 22B parameters per token for efficient performance. With a 256K context window and strong coding capabilities, it rivals top-tier proprietary models.
25 June 2026
Qwen3-30B-A3B: specs, benchmarks, and how to run it on Lyceum
Qwen3-30B-A3B activates only 3 billion parameters per token, delivering the reasoning capabilities of a 30B model at high speeds. Learn how to deploy this cost-efficient MoE model on Lyceum's EU-sovereign infrastructure.
22 June 2026
Nemotron-3-Nano-30B: specs, benchmarks, and how to run it on Lyceum
NVIDIA's Nemotron-3-Nano-30B-A3B combines a Mamba-Transformer architecture with a Mixture-of-Experts design to deliver top-tier reasoning at a fraction of the compute cost. Here is how to deploy it on Lyceum's EU-sovereign infrastructure.
21 June 2026
MiniCPM-V 4.5: specs, benchmarks, and how to run it on Lyceum
MiniCPM-V 4.5 scores 77.0 on OpenCompass in an efficient 8B package. With its novel 3D-Resampler, it compresses video tokens by 96x, making long-video understanding highly cost-effective.
21 June 2026
MiniMax-M2.5: specs, benchmarks, and how to run it on Lyceum
MiniMax-M2.5 delivers frontier-level coding performance at a fraction of the cost of proprietary models. Learn how to deploy this 230B parameter MoE model on Lyceum's serverless platform.
19 June 2026
Image Ultra: specs, benchmarks, and how to run it on Lyceum
Image Ultra delivers high-quality image generation in under one second. Designed for latency-sensitive applications, it offers a drop-in OpenAI-compatible API on EU-sovereign infrastructure.
18 June 2026
gpt-oss-120b: specs, benchmarks, and how to run it on Lyceum
gpt-oss-120b brings OpenAI's reasoning capabilities to the open-source ecosystem. With 117B parameters and a sparse MoE architecture, it delivers o4-mini-level performance while fitting on a single 80GB GPU.
18 June 2026
Hermes-4-405B: specs, benchmarks, and how to run it on Lyceum
Hermes-4-405B introduces a hybrid reasoning mode that balances fast responses with deep, think-tag deliberation. Now available on Lyceum's European infrastructure, it delivers strong math and coding performance without the censorship of proprietary models.
17 June 2026
GLM-5.1: specs, benchmarks, and how to run it on Lyceum
GLM-5.1 is a Mixture-of-Experts model with 754B parameters and 40B active per token, built for sustained, multi-step software engineering tasks. With a leading SWE-Bench Pro score among the models on its own card, it offers an open-weight alternative to frontier proprietary models.
16 June 2026
FLUX.1 Dev: specs, benchmarks, and how to run it on Lyceum
FLUX.1 Dev brings strong prompt adherence and photorealism to open-weights image generation. Learn how to deploy this 12B parameter rectified flow transformer on Lyceum's EU-hosted infrastructure.
16 June 2026
FLUX.2 Klein: specs, benchmarks, and how to run it on Lyceum
FLUX.2 Klein optimizes the speed-to-quality ratio for AI image generation. With a unified architecture for text-to-image and editing, it delivers photorealistic 1024x1024 outputs in under a second.
2 June 2026
Run Vision Language Models on GPU Cloud: VRAM & Setup Guide
Vision language models consume massive VRAM for image tokens. Learn the exact hardware requirements and deployment strategies for production VLMs.
31 May 2026
Multimodal AI Inference on European GPUs: Compliance and Cost Optimization
Running multimodal AI inference at scale exposes the structural flaws of hyperscaler pricing and compliance models. Engineering teams require infrastructure that provides high throughput for complex data types while maintaining strict data residency.
29 May 2026
Deploy Qwen 2.5 72B on GPU Cloud: VRAM Sizing and vLLM Setup
Running Qwen 2.5 72B in production requires strict memory management and the right infrastructure. Learn how to calculate VRAM requirements, configure vLLM, and deploy on EU-sovereign GPUs without hyperscaler price premiums.
28 May 2026
Deploy Gemma 3 on European GPU Cloud: VRAM, Setup, and GDPR Compliance
Google's Gemma 3 models bring multimodal capabilities and 128K context windows to open weights AI. Running them in production requires careful VRAM planning and infrastructure that guarantees data residency.
27 May 2026
Deploy DeepSeek R1 on European GPU Cloud: VRAM, Costs, and Compliance
Deploying DeepSeek R1 requires massive VRAM and strict data governance. Learn how to size your hardware and run production inference on EU-sovereign infrastructure without hyperscaler markups.