AI This article was created with the help of AI.
The same model at different prices
Enterprise engineering teams transitioning from proprietary per-seat coding tools to open-weight architectures frequently land on Qwen3-235B. The model card records 235 billion total parameters with 22 billion activated per forward pass, along with gains in instruction following, logical reasoning, mathematics and coding that Qwen reports for the 2507 release. However, when evaluating hosted inference endpoints, infrastructure leads immediately encounter a confusing pricing landscape: published rates for serving the exact same model weights differ by a wide margin from one provider to the next.
On public model aggregators such as OpenRouter, the listed endpoint for Qwen3-235B-A22B is priced at $0.455 per million input tokens and $1.82 per million output tokens, with a 131K-token context window, as shown on the model's live listing at the time of writing. Discounted or promotional endpoints elsewhere in the market sit well below that level, and extended thinking or reasoning modes are typically priced higher again. For an organisation routing tens of millions of prompt tokens daily across hundreds of developers, this variance translates into thousands of euros in monthly operational spend.
Headline figures alone do not explain these price differences. Treating two endpoints as identical simply because they advertise the same model identifier creates serious operational risk. When one host charges a third of the market rate for the same open weights, the difference usually stems from architectural compromises in execution, unstated quantization, constrained context windows, or offshore data routing. Evaluating per-token economics requires stripping away the headline rate and auditing the underlying infrastructure stack.
- Headline rate variance: OpenRouter lists this model as hosted by a single upstream provider at a 131K context window, and its per-token rates sit well above the cheapest hosted offers in the market.
- Hidden infrastructure trade-offs: Aggressive cost reduction often relies on unannounced weight quantization, aggressive KV-cache eviction, or offshore compute clusters.
- Enterprise evaluation requirements: Reliable procurement requires standardising precision, context length, concurrency, data residency, and cache discounting before comparing rates.
What the Lyceum record states
To establish a clear technical baseline, we publish verified hosting specifications and commercial terms directly in the model catalogue with per-model records. We operate under a strict region-honest policy where data sovereignty claims are tied directly to the physical deployment zone of each model rather than abstract corporate marketing.
For Qwen3-235B-A22B on Serverless Inference, the catalogue record confirms the following verified operational specifications:
- Model Identifier: Qwen/Qwen3-235B-A22B-Instruct-2507
- Deployment Region: eu-north1, physically hosted and executed within the European Union
- Data Governance: Zero data retention (ZDR) policy, no prompt logging for model training, and strict alignment with European data protection frameworks
- Input Token Pricing: $0.20 per 1M input tokens (exact rate, unrounded)
- Output Token Pricing: $0.60 per 1M output tokens (exact rate, unrounded)
- Billing Model: Pay-per-token serverless consumption with zero base fees, zero minimum commitments, and zero egress penalties
These metrics provide a stable anchor for enterprise cost modelling. Because the workload runs entirely within EU jurisdiction on dedicated infrastructure, procurement and compliance teams can evaluate total cost of ownership without factoring in legal risks or hidden transfer surcharges.
Five dimensions behind a price gap
When comparing Qwen3-235B pricing across providers, differences in the published rates trace back to five technical and operational dimensions. Without fixing these five parameters, direct price comparisons are invalid.
Served precision and weight quantization
Running Qwen3-235B at native 16-bit precision (BF16 or FP16) requires several hundred gigabytes of VRAM just to hold the model parameters, which in practice means a multi-GPU node: Qwen's own deployment guidance for the 2507 release launches the model with tensor parallelism across eight accelerators and warns that context length must be cut back if the node runs out of memory. Quantizing the weights to 8-bit (FP8) or 4-bit (AWQ/GPTQ) cuts the memory footprint substantially, enabling execution on smaller, cheaper GPU nodes. Providers running quantized weights operate at significantly lower hardware costs but often market the endpoint under the base model name without disclosing the precision loss.
Maximum supported context window
While the Qwen3-235B architecture supports long sequence processing, managing large context windows demands substantial GPU memory for the key-value (KV) cache. A provider offering a full 131K-token or 256K-token context window must reserve significant high-bandwidth memory per concurrent request, reducing overall batch concurrency. Providers advertising rock-bottom token rates frequently restrict maximum context to 8K or 16K tokens to maximize batch throughput on constrained hardware.
Sustained throughput and latency under concurrency
Per-token pricing does not guarantee service responsiveness. Low-cost providers often oversubscribe underlying GPU instances, leading to severe time-to-first-token (TTFT) degradation and erratic inter-token latency during peak enterprise hours. High-performance endpoints maintain strict memory headroom and continuous batching pipelines to sustain consistent token throughput across high concurrency.
Processing region and regulatory compliance
Data center operating costs, electricity tariffs, and compliance frameworks vary dramatically across geographical regions. Hosting inference workloads within the European Union under strict GDPR and EU AI Act standards carries distinct infrastructural costs compared to routing tokens through lower-cost facilities in jurisdictions subject to the US CLOUD Act or unverified foreign privacy standards.
Cached input token pricing
Modern inference engines leverage prefix caching to avoid recomputing attention states for repeated system prompts, few-shot examples, and shared document contexts. Providers that support prompt caching frequently offer substantial discounts on cached input tokens, significantly lowering the effective blended cost for agentic coding tools and iterative multi-turn dialogues. Checking live per-token pricing across providers reveals whether prompt caching is supported and how cache read hits are billed.
Precision is the most common hidden difference
Among the variables driving inference price disparities, served precision is the most frequent and least transparent factor. When two providers charge wildly divergent rates for Qwen3-235B, the lower rate often reflects an unannounced down-casting of the model weights.
To understand the cost mechanics, consider the VRAM math for a 235-billion parameter architecture. In uncompressed FP16 or BF16 format, each parameter occupies 2 bytes of GPU memory, so the baseline weights alone run to several hundred gigabytes, before any allocation for the KV cache and runtime activation buffers. That is why full-precision serving is spread across a multi-GPU node rather than a single card: Qwen's own deployment commands for the 2507 release launch the model with tensor parallelism across eight GPUs and note that context length must be reduced if the node hits out-of-memory errors. The hourly cost of operating that hardware tier defines the baseline floor for uncompressed inference.
In contrast, applying 4-bit weight quantization compresses each parameter to roughly half a byte, shrinking the static weight footprint to a small fraction of the full-precision figure and allowing the model to be served from a far smaller pool of accelerators than a native BF16 deployment needs. By slashing the physical hardware allocation, the host can sharply undercut market prices while preserving healthy operating margins.
Advanced inference engines like vLLM, whose serving stack is developed in the open, expose engine arguments that set weight precision and cache memory behaviour at launch: --dtype selects the data type used for model weights and activations, with half (FP16) recommended for AWQ quantization, and --kv-cache-dtype selects the data type for KV cache storage. While aggressive quantization delivers dramatic infrastructure savings, it introduces subtle precision degradation in complex reasoning, mathematical proofs, and syntactically strict code generation. Enterprise teams evaluating endpoints must verify whether a provider is serving uncompressed weights or quantized approximations.
Why a 235B model can be inexpensive per token
Engineers unaccustomed to Mixture-of-Experts architectures often wonder how a 235B model can be served at rates comparable to much smaller dense models. The answer lies in the structural decoupling of memory capacity and active compute.
In a traditional dense transformer, every forward pass activates the full parameter set for every token generated. Serving a 70B dense model requires computing 70 billion parameter interactions per token. In contrast, Qwen3-235B-A22B employs a sparse Mixture-of-Experts routing mechanism: while the complete model contains 235 billion parameters distributed across specialised expert layers, only 22 billion are activated for any given token.
This architectural split creates two distinct cost dynamics:
- Memory capacity is governed by total parameters: the system must provide enough VRAM, several hundred gigabytes in BF16, to retain all 235 billion weights in fast memory ready for routing.
- Compute cost is governed by active parameters: Floating-point operations (FLOPs) per token are calculated based purely on the 22 billion active parameters, keeping compute intensity and kernel execution times exceptionally lean.
Because the active compute footprint matches that of a compact 22B dense network, the operational compute cost per forward pass remains low. High-throughput serving depends just as much on memory management: the PagedAttention paper behind vLLM reports that KV-cache memory is otherwise significantly wasted by fragmentation and redundant duplication, which limits batch size, and that its paging approach achieves near-zero waste and improves throughput by 2-4 times at the same latency compared with earlier serving systems. That throughput is what lets a host quote a low per-token rate. Managing the balance between context length and GPU memory requirements ensures that the host can maintain dense batching without triggering out-of-memory faults during peak loads.
Detailed architectural specifications are published in the model card, which records 94 layers, grouped-query attention with 64 query heads and 4 key-value heads, and 128 experts with 8 activated per token. Hugging Face documents the model card as the README file that accompanies every model repository and describes the model, its intended uses and limitations, training parameters and evaluation results, so engineers can inspect exactly how the sparse layers are wired before committing to a provider.
Comparing only providers that match your requirements
Declaring a single provider as the cheapest option across the board is a flawed evaluation strategy. A provider offering the lowest headline rate on a 4-bit quantized model with an 8K context limit is entirely unsuitable for a multi-file enterprise coding agent requiring 64K context and native FP16 precision. Enterprise buyers must apply a standardized, like-for-like evaluation framework.
Standardised benchmarking frameworks, such as the MLPerf Inference suite, establish rigorous division rules to prevent misleading comparisons. MLCommons states that the Closed division is intended to compare hardware platforms or software frameworks apples-to-apples and requires using the same model as the reference implementation, while the Open division allows using a different model or retraining. The same discipline applies to open-weight deployments: the Open Source Initiative's definition applies its requirements equally to a system, a model, and its weights and parameters, so a like-for-like comparison should confirm which weights a provider is actually running.
To conduct an objective cross-provider comparison, follow a structured three-step qualification process:
- Fix your non-negotiable operational constraints: Define your required weight precision (native BF16 vs FP8/INT4), minimum concurrent context window (e.g., 32K or 128K tokens), target time-to-first-token latency, and mandatory data residency jurisdiction.
- Filter out mismatched providers: Disqualify any vendor that fails to meet your technical or regulatory floor, regardless of how aggressively their headline rates undercut the market.
- Compare effective blended cost across qualified candidates: Evaluate pricing using your real-world prompt-to-completion token ratios, accounting for cached input discounts, tool-calling overhead, and output generation rates.
For developer-facing workloads involving structured output and automated tool orchestration, verifying provider compatibility with open models with reliable function calling ensures that API endpoints deliver predictable JSON schemas without parse failures.
Deploying Qwen3-235B-A22B for EU workloads
For European enterprises seeking to replace expensive proprietary closed-model seat licenses with open-weight infrastructure, Qwen3-235B-A22B delivers the reasoning depth, instruction-following accuracy, and multilingual versatility required for demanding production workflows.
At Lyceum, we deliver Qwen3-235B-A22B through our Serverless Inference platform, hosted entirely in eu-north1 on owned European infrastructure. We serve the model with transparent precision standards, zero data retention, and strict compliance with European privacy standards, ensuring complete isolation for sensitive corporate codebases and customer data.
- Per-token billing: input and output are metered and priced separately, not as one blended rate
- Input rate: $0.20 per 1M input tokens (exact, unrounded)
- Output rate: $0.60 per 1M output tokens (exact, unrounded)
- Hosting zone: eu-north1 (European Union)
- Integration: Fully OpenAI-compatible chat completions endpoint with zero minimum spend
When evaluating infrastructure for enterprise AI deployments, fix the precision, context, and data residency requirements your workloads demand, then select the platform that delivers verified execution on those terms. Teams migrating production workloads can explore GDPR-compliant LLM inference in Europe to establish a defensible, high-performance foundation for open-weight models.