Text

17 articles

Articles

31 July 2026

Kimi K3 API: Where to Run It, and What a 1M-Token Context Costs

Kimi K3 is Moonshot AI's open-weights model with a 1M-token context window. On Lyceum it runs as moonshotai/kimi-k3 at $3.00 input, $0.75 cached input and $15.00 output per million tokens, EU-hosted, with prompts and outputs not stored or used for training. Kimi K2.6 was retired on 5 October 2026, and its model id now routes to Kimi K3.

8 September 2026

GLM-5.2 Instant: Latency-Optimised GLM, EU-Hosted

GLM-5.2 Instant offers the same 1M-token context window and identical per-token pricing as the base GLM-5.2 model, but is positioned for lower latency. It runs on EU-sovereign serverless inference in eu-north1, so teams can evaluate throughput on their own traffic.

8 September 2026

Qwen3.5-9B: Compact Qwen With the Narrowest Price Spread

Qwen3.5-9B is a compact open-weight model with a 256K context window and the narrowest input-to-output price spread in the catalogue. Costing $0.15 per million input tokens and $0.20 per million output tokens, it significantly reduces the total bill for output-heavy workloads.

12 August 2026

Where to Run Kimi Models in Europe: K2.6, K2.7 Code and K3

Moonshot AI's Kimi models deliver frontier capabilities for agentic coding. K2.6 and K2.7 Code offer 1T-parameter scale with 256K context, while K3 pushes to 2.8T parameters and a 1M-token window. European teams can run them via EU-hosted APIs to maintain data residency.

27 June 2026

GLM-5.2: specs, benchmarks, and how to run it on Lyceum

7 September 2026

MiniMax-M3: EU-Hosted Where M2.5 Is Global

If you are evaluating MiniMax under a data-residency requirement, this page states which of the two models has a confirmed European region and what choosing it costs you per token.

20 August 2026

DeepSeek-V4-Flash: specs, benchmarks, and how to run it

DeepSeek-V4-Flash is a 284-billion parameter MoE model offering agentic reasoning across a 1-million token context window. Lyceum serves it via an OpenAI-compatible API from eu-north1 in the European Union, optimized for enterprise inference at $0.15 per million input tokens.

1 August 2026

DeepSeek V4 Flash: 1M-Token Context for AI Products

DeepSeek V4 Flash introduces a 284B parameter MoE architecture with 13B active parameters, delivering low time-to-first-token latency and a 1,048,576-token context window. For AI-native products, this means high-throughput agent loops and long-context retrieval hosted natively in Europe

16 May 2026

Deploying Microsoft Phi-4 Inference on GPU Cloud: A Production Guide

25 June 2026

Qwen3-235B-A22B: specs, benchmarks, and how to run it on Lyceum

Qwen3-235B-A22B-Instruct-2507 is Alibaba's flagship Mixture-of-Experts model, activating only 22B parameters per token for efficient performance. With a 256K context window and strong coding capabilities, it rivals top-tier proprietary models.

25 June 2026

Qwen3-30B-A3B: specs, benchmarks, and how to run it on Lyceum

Qwen3-30B-A3B activates only 3 billion parameters per token, delivering the reasoning capabilities of a 30B model at high speeds. Learn how to deploy this cost-efficient MoE model on Lyceum's EU-sovereign infrastructure.

22 June 2026

Nemotron-3-Nano-30B: specs, benchmarks, and how to run it on Lyceum

NVIDIA's Nemotron-3-Nano-30B-A3B combines a Mamba-Transformer architecture with a Mixture-of-Experts design to deliver top-tier reasoning at a fraction of the compute cost. Here is how to deploy it on Lyceum's EU-sovereign infrastructure.

18 June 2026

gpt-oss-120b: specs, benchmarks, and how to run it on Lyceum

gpt-oss-120b brings OpenAI's reasoning capabilities to the open-source ecosystem. With 117B parameters and a sparse MoE architecture, it delivers o4-mini-level performance while fitting on a single 80GB GPU.

18 June 2026

Hermes-4-405B: specs, benchmarks, and how to run it on Lyceum

Hermes-4-405B introduces a hybrid reasoning mode that balances fast responses with deep, think-tag deliberation. Now available on Lyceum's European infrastructure, it delivers strong math and coding performance without the censorship of proprietary models.

17 June 2026

GLM-5.1: specs, benchmarks, and how to run it on Lyceum

GLM-5.1 is a Mixture-of-Experts model with 754B parameters and 40B active per token, built for sustained, multi-step software engineering tasks. With a leading SWE-Bench Pro score among the models on its own card, it offers an open-weight alternative to frontier proprietary models.

29 May 2026

Deploy Qwen 2.5 72B on GPU Cloud: VRAM Sizing and vLLM Setup

Running Qwen 2.5 72B in production requires strict memory management and the right infrastructure. Learn how to calculate VRAM requirements, configure vLLM, and deploy on EU-sovereign GPUs without hyperscaler price premiums.

28 May 2026

Deploy Gemma 3 on European GPU Cloud: VRAM, Setup, and GDPR Compliance

Google's Gemma 3 models bring multimodal capabilities and 128K context windows to open weights AI. Running them in production requires careful VRAM planning and infrastructure that guarantees data residency.

Your next workload starts here