EU inference providers compared: how this table was built

Choosing an inference endpoint requires more than a low token price. This comparison covers selected providers that buyers consider for European workloads, plus a price comparator whose processing location needs confirmation. It is a shortlist, not an exhaustive market survey.

The table in the next section follows five rules, applied to every provider including the one publishing this article:

  • Each row names the provider, model, processing scope, price and compatibility
  • Prices use the vendor’s published currency and were checked on 1 October 2026
  • An unverified processing location means further evidence is needed, not that a location is unavailable
  • Only identical named models support a direct price comparison
  • Recheck prices and contract terms before buying

An EU processing requirement may come from a customer contract, internal policy or a legal assessment. GDPR does not automatically prohibit every US-hosted service. For personal data, assess the processing terms and any qualifying third-country transfer. Lyceum publishes this comparison and competes with the providers listed.

The providers serving open models from Europe

An open-weight model can be served by several providers. Matching its name does not prove identical weights, quantisation, chat templates or runtime settings. These differences can affect quality, speed and usage accounting, so test the actual endpoint.

Mistral offers regional endpoints, OVHcloud offers a managed open-model catalogue, and Lyceum states hosting per model. DeepInfra is included as a price comparator. This review has not established an EU processing commitment for the DeepInfra endpoint, so request one if your workload requires it.

OVHcloud’s gpt-oss-120b page documents an OpenAI-compatible endpoint and 400 authenticated requests per minute, per Public Cloud project and per model. A service’s worldwide customer availability is not evidence that inference leaves Europe. Verify the processing location and contract separately.

Models, regions and per-token prices compared

Prices below were checked on 1 October 2026. Amounts are per million tokens, in the listed currency. Different models are examples of each catalogue, not quality-equivalent alternatives. Mistral’s regional rate is calculated from its standard price and documented 10% premium; confirm model availability on that endpoint.

ProviderReference modelProcessing regionInput / output per 1M, 1 October 2026OpenAI compatibilityRetention check
LyceumGLM-5.3; DeepSeek-V4-Flash-0731Both listed EU-hosted in dashboardGLM-5.3: $1.40 / $4.40; DeepSeek: $0.25 / $0.30 (USD)Documented chat completions; separate Anthropic routeSelf-asserted zero retention for inference; DPA on request
MistralMistral Large 3EU endpoint covers EU and EFTA; global control planeStandard $0.50 / $1.50; regional $0.55 / $1.65 (USD), subject to model availabilityDocumented chat-completions APICheck eligible endpoints, zero-retention settings and product exceptions
OVHcloudgpt-oss-120bConfirm contracted AI Endpoints processing location€0.08 / €0.40 (EUR)Documented, with streaming and function callingConfirm prompt, output, billing and log treatment in applicable terms
DeepInfraDeepSeek-V4-Flash-0731EU commitment not established in this review$0.06 / $0.18 (USD)Documented OpenAI-compatible chat APICheck model-specific, batch and third-party processing exceptions

DeepInfra’s listed DeepSeek-V4-Flash-0731 price is lower than Lyceum’s matched model price. That comparison does not settle quality, latency, capacity or residency. The Mistral and OVHcloud rows use different models, so they cannot establish which provider is cheapest for an equivalent result.

Lyceum’s dashboard lists the 2 models in its row as EU-hosted on 1 October 2026. Compare total cost per successful task, including cache misses, retries and reasoning output. At sustained volume, also compare dedicated capacity with the token bill.

OpenAI compatibility and what breaks on a swap

All 4 providers document an OpenAI-compatible chat interface. Configure the provider’s base URL, key and model. Test the specific request fields and responses your application needs; broad compatibility does not establish feature parity.

  • Tool calling. Schema interpretation varies across open weights. A tool definition your closed model honoured verbatim can be ignored, hallucinated or wrapped differently by an open model served behind the same endpoint, and agent loops fail on the third call, not the first.
  • Streaming shape. Chunk structure and finish-reason fields differ across serving stacks. Clients that parse deltas strictly, or that branch on finish_reason, need a regression pass against the new endpoint's stream.
  • Token accounting. Usage fields and tokenizer differences change the cost math. Two providers serving the same weights can report different token counts for the same prompt, which moves your per-seat cost model in either direction.

Use the actual endpoint as the test target. A serving engine name alone does not establish compatible tool calls or streaming behaviour. Record the model, settings and client version so a later change can be compared.

Retention and training policies compared

Article 28 of the GDPR sets the contract your processor has to meet: processing only on documented instructions from the controller, authorisation before another processor is engaged, and deletion or return of all personal data at the end of the service. Retention policy is where that legal frame meets the provider's architecture, and it is the column procurement reads most closely.

Review retention at the product and endpoint level. Ask separately about prompts, outputs, application logs, abuse monitoring, billing records and backups. Check whether zero retention is automatic or requires approval, and whether batch jobs, preview models or third-party models have exceptions. A no-training statement does not by itself mean no storage.

Lyceum self-asserts zero data retention for inference: prompts and outputs are not stored in a database or used for training. Short-lived per-session caching can remain in GPU memory. A DPA is available on request. Confirm the scope and supporting terms during review; this is not a certification.

Where the published data is thin or missing

Four columns are thin across the market, and knowing which ones tells you what to ask each vendor:

  • Processing scope: obtain a commitment covering the exact endpoint, model, failover and support access. A valid service-wide commitment can also cover individual models
  • Retention: distinguish request content from operational metadata and confirm exceptions
  • Training: check the applicable product terms and account settings
  • Capacity: request account-specific rate limits and test realistic concurrency. OVHcloud documents 400 requests per minute per project and model for the reference endpoint

A dated 'not published' is more useful to procurement than a guess, because it names the exact question to put to the vendor: which certification, which region, which retention window, asked in writing and answered with a date.

Lyceum publishes no service-level agreement, uptime target or service credit for serverless inference. If you require contractual guarantees, obtain a suitable written offer before selecting the service.

Which provider fits which workload

Map your situation to the column that decides it, and the shortlist writes itself:

  • Lower listed token cost: compare DeepInfra’s matched DeepSeek price, then measure quality, latency and total task cost
  • European processing: confirm the selected Mistral, OVHcloud or Lyceum endpoint against your actual location and retention requirements
  • Coding agents: test tool calls, streaming, reasoning fields and retries. Lyceum documents an Anthropic-format route for Claude Code
  • Procurement: prefer complete, dated evidence for region, retention, subprocessors and account limits

Three takeaways:

  • Verify the processing commitment for the model and product you will use
  • A lower token rate is one input to total cost per accepted result
  • Run representative compatibility and quality tests before switching production traffic

Use the current Lyceum dashboard and each competitor’s pricing page when making a purchase decision. This comparison is a dated snapshot and does not promise a future refresh schedule.