AI This article was created with the help of AI.

The Economics of Replacing Closed APIs

When engineering teams look for an open source alternative to openai api endpoints or Anthropic models, the immediate driver is usually financial waste. Routing every internal classification task, customer support turn, and batch extraction script through a frontier closed model creates unsustainable cloud expenditure. A closed-API bill is one blended price applied to jobs that are not alike, so unbundling homogeneous API calls into workload-specific open models is what actually moves the metered spend. Treating a single proprietary model as a universal backend forces you to pay frontier rates for tasks that require only a fraction of that parameter capacity.

The mistake many engineering organizations make during this transition is searching for a 1:1 drop-in model replacement. There is no single open-weight checkpoint that matches every nuance of a proprietary system across reasoning, coding, classification, and multi-turn dialogue. Instead, replacing a closed API key requires a portfolio strategy: mapping specific production workloads to specialized open models optimized for those exact compute profiles. By separating model capability from token volume, teams route high-frequency, structured tasks to lightweight weights while reserving massive mixture-of-experts (MoE) architectures for complex reasoning.

From an architectural standpoint, this migration functions strictly as a model-routing decision. When moving to managed Serverless Inference, you do not need to solve the complex capacity planning and utilization crossover math of self-hosting dedicated GPU clusters. You retain the consumption flexibility of per-token metering while eliminating the margin premiums of closed providers.

Everyday Chat and High-Volume Extraction

Routine conversational interfaces, internal assistants, and customer support pipelines represent a massive share of production token traffic. For these multi-turn interactions, standard dense models provide predictable token cadence and low time-to-first-token (TTFT). Three core open weights serve this category directly from the eu-north1 region:

  • Llama-3.3-70B (API string: meta-llama/Llama-3.3-70B-Instruct): Priced at $0.13 per 1M input tokens and $0.40 per 1M output tokens in the Fast tier. Ideal for complex multi-turn conversational agents that require broad world knowledge and robust instruction following.
  • Hermes-4-70B (API string: NousResearch/Hermes-4-70B): Priced at $0.13 per 1M input tokens and $0.40 per 1M output tokens in the Fast tier. Specifically tuned for natural conversation, prompt alignment, and multi-turn steering.
  • Gemma-3-27B (API string: google/gemma-3-27b-it): Priced at $0.10 per 1M input tokens and $0.30 per 1M output tokens. An instruction-tuned 27B parameter dense model that delivers low-latency responses for standard agent workflows.

High-Throughput Extraction and Classification Pipelines

For high-volume structured workloads like document classification, entity tagging, JSON schema extraction, and batch summarization, paying for large parameter footprints creates severe cost drag. Highly sparse Mixture-of-Experts (MoE) architectures and compact dense models handle these structured schemas with minimal latency, supported by modern serving runtimes that split key and value cache data into fixed-size blocks so attention memory can be paged rather than pre-reserved.

ModelAPI Model StringInput Price (/1M)Output Price (/1M)Hosting RegionPrimary Workload
Nemotron-3-Nano-30Bnvidia/NVIDIA-Nemotron-3-Nano-30B-A3B$0.06$0.24eu-north1High-volume extraction, compact MoE
Qwen3-30B-A3BQwen/Qwen3-30B-A3B-Instruct-2507$0.10$0.30eu-north1Instruction following, structured JSON
Qwen3-32BQwen/Qwen3-32B$0.10$0.30eu-north1Dense reasoning and tagging
Cosmos3-Super-Reasonernvidia/Cosmos3-Super-Reasoner$0.10$0.30eu-north1Complex multi-step extraction
Qwen3.5-9B(No published string)$0.15$0.20No published regionLong output extraction (256K context)

When selecting extraction models, evaluate input-to-output ratios. Qwen3.5-9B features a 256K-token context window with a $0.15 input and $0.20 output structure. This narrow spread makes it efficient for tasks where generated output length approaches input length, such as detailed report reformatting. Note that Qwen3.5-9B carries no publishable hosting region and no published API model string; verify regional constraints before routing regulated payloads to it.

Heavy Reasoning and Agentic Workloads

Workloads involving multi-step planning, scientific analysis, mathematical proof generation, and autonomous agent loops demand deep parameter capacity. When replacing proprietary reasoning endpoints, engineering teams must weigh raw parameter scale against latency and token cost. Below is the mapping for heavy reasoning models available across regions:

  • gpt-oss-120b (API string: openai/gpt-oss-120b): $0.15 per 1M input tokens, $0.60 per 1M output tokens, hosted in eu-north1. An open-weight 120B model catalogued for reasoning work on European infrastructure.
  • Qwen3-235B-A22B (API string: Qwen/Qwen3-235B-A22B-Instruct-2507): $0.20 per 1M input tokens, $0.60 per 1M output tokens, hosted in eu-north1, with a 256K-token context window. A large MoE catalogued for reasoning and instruction following.
  • Nemotron-Ultra-253B (API string: nvidia/Llama-3_1-Nemotron-Ultra-253B-v1): $0.60 per 1M input tokens, $1.80 per 1M output tokens, hosted in eu-north1. Built on Llama-3.1 foundations for multi-hop reasoning.
  • Hermes-4-405B (API string: NousResearch/Hermes-4-405B): $1.00 per 1M input tokens, $3.00 per 1M output tokens, hosted in eu-north1. A 405B parameter dense model for high-complexity analytical tasks.
  • Qwen3.5-397B-A17B (API string: Qwen/Qwen3.5-397B-A17B): $0.60 per 1M input tokens, $3.60 per 1M output tokens, hosted Global (multi-region). This model is not pinned to the European Union, so no EU, GDPR or data-residency claim applies to it.

Autonomous Agents and Tool Execution

Agentic loops require models that reliably handle schema enforcement, function calling, state persistence, and error recovery across long horizons. GLM-5.2 (API string: zai-org/GLM-5.2) serves as the flagship EU-hosted option in eu-north1 at $1.50 per 1M input tokens and $4.50 per 1M output tokens, delivering bilingual reasoning, 1M-token context support, and stable tool invocation. For lightweight edge agents, Nemotron-3-Nano-Omni (API string: nvidia/Nemotron-3-Nano-Omni) runs in eu-north1 at $0.06 per 1M input tokens and $0.24 per 1M output tokens.

For specialized agentic capabilities, DeepSeek-V4-Pro (API string: deepseek/deepseek-v4-pro-0813) is priced at $2.00 per 1M input tokens and $4.00 per 1M output tokens, and Kimi-K2.6 (API string: moonshotai/Kimi-K2.6) is priced at $1.00 per 1M input tokens and $4.00 per 1M output tokens. Both are catalogued for long-horizon agentic workflows and multimodal context. We state no publishable hosting region for DeepSeek-V4-Pro or Kimi-K2.6, and make no EU sovereignty, GDPR residency, or data-protection claims regarding their endpoints. Similarly, Nemotron-3-Ultra-550b (API string: nvidia/Nemotron-3-Ultra-550b-a55b) is priced at $1.00 per 1M input tokens and $3.00 per 1M output tokens and is Global (multi-region), not pinned to the EU.

Code Generation and Code Review

Migrating automated software engineering pipelines, pull request review bots, and inline code completion engines away from proprietary endpoints starts with what the catalogue actually publishes: price, context window, and whether a region is on record. No first-party benchmark or capability score is published for these models, so treat the descriptions as positioning and run your own evaluation on your own repositories before committing a pipeline.

Specialized code models currently offered include:

  • Qwen3-Coder-30B-A3B: Billed at $0.06 per 1M input tokens and $0.25 per 1M output tokens with a 256K-token context window. This 30B MoE activates only 3B parameters per token, tying the platform minimum for text input costs. It publishes no API model string.
  • Kimi-K2.7-Code: Billed at $1.25 per 1M input tokens and $4.50 per 1M output tokens with a 256K-token context window. A 1T parameter MoE built for agentic code refactoring and long-context repo parsing. It publishes no API model string.
  • DeepSeek-V4-Pro (API string: deepseek/deepseek-v4-pro-0813): Billed at $2.00 per 1M input tokens and $4.00 per 1M output tokens, catalogued specifically for complex software engineering and multi-file code review.

An essential architectural boundary must be stated clearly: no code-specialized model in the catalogue carries a verified EU hosting record. DeepSeek-V4-Pro, Qwen3-Coder-30B-A3B, and Kimi-K2.7-Code have no publishable region, and no data sovereignty or GDPR residency claim applies to them. Engineering teams bound by strict European data residency requirements must route code generation workloads through general EU-hosted models in eu-north1, such as GLM-5.2 ($1.50 in / $4.50 out), Qwen3-235B-A22B ($0.20 in / $0.60 out), or Cosmos3-Super-Reasoner ($0.10 in / $0.30 out).

Large Context Windows and Long Documents

Legal discovery, full-codebase indexing, financial audit trails, and multi-document synthesis require context windows extending up to 1 million tokens. When processing massive context windows, memory held in the Key-Value (KV) cache grows with sequence length, which is why serving runtimes rely on paged memory allocation to avoid GPU out-of-memory errors.

ModelContext WindowInput Price (/1M)Output Price (/1M)Published RegionAPI String Published
Kimi-K31,000,000$3.00$15.00No published regionNo
DeepSeek-V4-Flash1,000,000$0.15$0.30No published regionNo
MiniMax-M31,000,000$0.40$2.00No published regionNo
GLM-5.2 Instant1,000,000$1.50$4.50No published regionNo (unverified)
GLM-5.21,000,000$1.50$4.50eu-north1Yes (zai-org/GLM-5.2)
Hermes-4-405B128,000$1.00$3.00eu-north1Yes (NousResearch/Hermes-4-405B)

Four prominent 1M-token models handle ultra-long documents on the platform: Kimi-K3 ($3.00 in / $15.00 out, a 2.8T parameter MoE holding the highest per-token rate in the catalogue), DeepSeek-V4-Flash ($0.15 in / $0.30 out, a latency-optimized variant at about one-thirteenth the cost of Pro), MiniMax-M3 ($0.40 in / $2.00 out), and GLM-5.2 Instant ($1.50 in / $4.50 out). None of these four models publishes an API model string, and none carries a confirmed hosting region.

For teams requiring verified EU hosting for large documents, GLM-5.2 ($1.50 in / $4.50 out) provides a full 1M-token context in eu-north1. Hermes-4-405B ($1.00 in / $3.00 out) operates in eu-north1 with a 128K context window. Kimi-K2.6 ($1.00 in / $4.00 out) offers a 256K window without a published region. Context windows reflect published dashboard specifications as of August 2026; always verify effective context windows against upstream model documentation before sizing production prompts.

Multimodal Vision, Images, and Embeddings

Enterprise AI stacks frequently require multimodal capabilities alongside text inference, including vector retrieval for RAG architectures, document image parsing, and visual asset generation. These modalities are served natively with per-token, per-image, or per-second billing.

Retrieval-Augmented Generation (RAG) pipelines rely on dense vector embeddings to ground generation in private documents. Qwen3-Embedding-8B (API string: Qwen/Qwen3-Embedding-8B) is hosted in eu-north1 on the /embeddings endpoint at $0.01 per 1M tokens. It generates 4,096-dimensional dense multilingual embeddings over a 32K-token context, covering semantic retrieval across multi-language enterprise knowledge bases.

Vision and Document Understanding

For visual question answering, OCR extraction, chart analysis, and diagram interpretation, two open vision-language models run in eu-north1:

  • Qwen2.5-VL-72B (API string: Qwen/Qwen2.5-VL-72B-Instruct): Priced at $0.25 per 1M input tokens and $0.75 per 1M output tokens. Supports combined image and text inputs for complex spatial reasoning, dense table extraction, and technical document comprehension.
  • MiniCPM-V 4.5 (API string: openbmb/MiniCPM-V-4_5): Priced at $0.66 per 1M input tokens and $1.11 per 1M output tokens. An efficient vision-language model optimized for fast document parsing and image QA.

Image Generation and Video Replacement

Image generation pipelines replace proprietary diffusion endpoints with open weights hosted in eu-north1. FLUX.2 Klein (API string: lyc-flux-2-klein) is metered at $0.0001 per image. Wan Image (API string: lyc-wan-image) and FLUX.1 Dev (API string: lyc-flux-1-dev) are billed at $0.005 per image. Image Ultra (API string: lyc-image-ultra) is billed at $0.005 per image and is catalogued as returning results in under one second. Video Replace is billed at $0.03 per second of output for 720p and $0.06 per second for 1080p; it performs subject and face replacement in existing video streams, is hosted in eu-north1, and publishes no API model string.

Executing the Switch: Serverless Inference

Migrating production pipelines from closed APIs to open weights does not require rewriting your application layer. Serverless Inference is fully OpenAI SDK compatible, so transitioning an OpenAI-SDK codebase requires updating only two parameters: the base URL of the serverless endpoint and the model string. An Anthropic-SDK codebase changes client library as well, because the endpoint speaks the OpenAI chat-completions shape. The underlying infrastructure runs on an open stack of vLLM, NVIDIA Dynamo, and TensorRT-LLM, with vLLM's paged attention implementation storing key and value cache data in separate blocks rather than one contiguous reservation.

For engineering teams running complex application topologies that cannot manually split traffic across individual endpoints, we provide four Smart Routing configurations:

  • Simple: routes basic conversational turns and extraction tasks to cost-efficient lightweight models.
  • Router: analyzes prompt complexity and dynamically dispatches requests to appropriate model tiers.
  • Reasoning: dispatches multi-step mathematical, analytical, and logical prompts to high-capacity reasoning architectures.
  • Complex: routes unstructured, long-horizon tasks to top-tier agentic models.

The dashboard publishes no fixed prices and no regional hosting records for these four Smart Routing entries. Consequently, we make no regional pinning or EU data-residency claims for routed traffic.

From a commercial and operational standpoint, Serverless Inference is a self-serve, pay-per-token product. It carries no SLA, no uptime target, no availability tier, and no service credits. Teams requiring contractual SLAs, private dedicated nodes, or custom containerized models can deploy Dedicated Inference, which provisions isolated endpoints across European data centres with automated scaling and scale-to-zero capabilities. For teams replacing closed endpoints with flexible, per-token billing, Serverless Inference delivers the open-model portfolio you need without infrastructure lock-in.