Qwen models processed in the EU

The Lyceum dashboard lists the 6 Qwen models below as EU-hosted on 1 October 2026. The first table covers 2 chat models and an embedding model; the next covers the Qwen3.8 line-up. Hosting statements apply to those specific offerings. Recheck the model record and processing terms before relying on residency.

ModelAPI stringRegionInput /1MCached input /1MOutput /1M
Qwen3-235B-A22Bqwen/qwen3-235b-a22b-instruct-2507EU, checked 1 October 2026$0.20$0.20$0.60
Qwen3.5-9Bqwen/qwen3.5-9bEU, checked 1 October 2026$0.15$0.04$0.20
Qwen3-Embedding-8Bqwen/qwen3-embedding-8bEU, checked 1 October 2026$0.02--

Qwen3-235B-A22B-Instruct-2507 has 235 billion total and 22 billion active parameters. This specific checkpoint supports non-thinking mode only. Qwen3.5-9B is a smaller chat and vision option with a 256K context. Qwen3-Embedding-8B produces vectors with up to 4,096 dimensions and a 32K context. Check the API route and supported input type for each.

Prices are USD per million tokens, checked on 1 October 2026. Qwen3-Embedding-8B lists at $0.02 per million input tokens. Recheck the live dashboard for current prices and the API model list for valid identifiers.

What the Qwen3 family brings: modes, licence and languages

The original Qwen3 family included dense and mixture-of-experts models, with thinking and non-thinking capabilities that varied by release. Do not apply that family description to every later checkpoint. The official Qwen3-235B-A22B-Instruct-2507 card explicitly states non-thinking-only operation and an Apache 2.0 licence.

  • Mode: verify the exact checkpoint, not just the Qwen family name
  • Licence: read the model card and licence for each release before downloading or modifying weights
  • Context: use the hosted endpoint’s supported window, which may differ from an upstream extension method
  • Evaluation: separate the model maker’s benchmark results from measurements on your workload

For retrieval workloads, the Qwen3 Embedding paper describes a separate series in 0.6B, 4B and 8B sizes for embedding and reranking, trained with a multi-stage pipeline on top of the Qwen3 backbones. Qwen reports state-of-the-art results on the multilingual MTEB benchmark and on code, cross-lingual and multilingual retrieval tasks, and the embedding models are also Apache 2.0. The 8B embedding model is the size served on the endpoint, as qwen/qwen3-embedding-8b on the embeddings path.

The Qwen3.8 line-up: sizes, context and prices

Lyceum lists 3 Qwen3.8 models with 256K context windows. The dashboard marks 27B and Flash-Next as vision-capable; the 2.4T-A95B entry is text-only. These are properties of the hosted offers checked on 1 October 2026. Check each upstream model card separately for architecture and licence terms.

ModelAPI stringHosted input supportContextInput /1MCached /1MOutput /1M
Qwen3.8-2.4T-A95Bqwen/qwen3.8-2.4t-a95bText256K$2.50$0.63$6.00
Qwen3.8-27Bqwen/qwen3.8-27bText and vision256K$0.40$0.10$2.40
Qwen3.8-Flash-Nextqwen/qwen3.8-flash-nextText and vision256K$0.20$0.05$0.50

For Qwen3.8-2.4T-A95B, a cache hit costs $0.63 per million input tokens versus $2.50 for fresh input, about 75% less for those cached tokens. An identical prefix can improve reuse, but caching is best effort. Budget fresh-input rates unless you have measured a stable hit rate.

Reasoning controls depend on the exact model and API identifier. Check the current model details before setting reasoning_effort or using an instant variant. The serverless API caps max_tokens at 65,536; this is an upper limit, not a recommended default. Choose a smaller output budget suitable for the task and check whether reasoning consumes it.

Choosing a Qwen model by the job

The right pick is a function of the job and the cost per million tokens, not of benchmark rank. Map the workload first, then read the price off that model's record.

Qwen3.5-9B lists at $0.15 input and $0.20 output per million tokens. Those rates alone cannot establish savings against a seat licence, which may bundle tools and different usage limits. Compare the full bill and accepted-task quality using a sample of your own traffic.

Region is a per-model record, not a family trait

Processing location matters because GDPR Article 44 sets the general principle that personal data undergoing processing may move to a third country only under the conditions laid down in Chapter V of the Regulation. The EDPB's guidance identifies three cumulative criteria for when a transfer outside the EEA occurs: a controller or processor is subject to the GDPR for the given processing; that controller or processor discloses by transmission or otherwise makes personal data available to another organisation; and that organisation is in a country outside the EEA or is an international organisation. An inference call that sends prompt data to infrastructure outside the EEA can meet all three.

Transfers are not prohibited, but they require a mechanism: an adequacy decision, or appropriate safeguards such as standard contractual clauses. The Commission's adequacy list covers countries including Japan, the United Kingdom and the United States for commercial organisations participating in the EU-US Data Privacy Framework, and exporters are responsible for checking that a decision is still in force. For a deeper treatment of what sovereignty does and does not mean, see GDPR-compliant LLM inference in Europe.

Check the model record, endpoint and failover commitment together. A dashboard EU label is a useful starting point, but does not establish compliance for your complete data flow. Include application logs, tools, support access and processing terms in the review.

What to verify before you switch

  • Context and output: the API documents a 65,536 max_tokens ceiling and a 300-second generation limit. Streaming does not remove those limits
  • Caching: treat cache hits as best effort and parse missing cache-usage fields defensively
  • Capacity: run representative concurrency tests and inspect current model notices before promising latency or throughput
  • Errors: confirm the current model identifier and account access. Use bounded retries only for documented transient failures

A successful request checks basic access. It does not establish quality, capacity or legal suitability. Run representative tests and record the model, settings, usage, errors and latency.

Moving a workload onto an EU-processed Qwen model

Configure 3 values in the OpenAI client: base URL, Lyceum API key and model identifier. This minimal chat example illustrates the request shape. Test streaming, tools and usage fields separately if your application uses them.

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.lyceum.technology/openai/v1",
    api_key=os.environ["LYCEUM_API_KEY"],
    timeout=30.0,
    max_retries=0,
)
resp = client.chat.completions.create(
    model="qwen/qwen3-235b-a22b-instruct-2507",
    messages=[{"role": "user", "content": "Summarise this sample: payment is due in 30 days."}],
    max_tokens=1024,
)
print(resp.choices[0].message.content)

Inference is billed per token with no minimum commitment. Test requests still incur token charges. Set a bounded evaluation budget and check the API documentation for request fields, reasoning controls and errors.

Qwen3-235B-A22B-Instruct-2507 is one candidate for text tasks at $0.20 input and $0.60 output per million tokens. It is non-thinking only. Choose between it, smaller chat and vision models, and the embedding endpoint using your own task results rather than the family name.