If a customer requires European processing, compare the exact Together AI product you use with alternatives that meet that requirement. Together offers dedicated deployments with regional choices. That does not establish where a serverless request runs. Ask for the region and failover policy for your model and tier before deciding that a migration is necessary.
GDPR Chapter V governs qualifying transfers to third countries. The EDPB identifies three cumulative conditions: an exporter subject to the GDPR makes data available to a separate controller or processor in a third country or international organisation. A foreign parent company alone does not establish such a transfer. Where a transfer exists, check an applicable adequacy decision, appropriate safeguards or a narrowly applicable derogation.
- An adequacy decision for the destination country, which puts transfers on the same footing as intra-EU flows.
- Appropriate safeguards, such as the standard data protection clauses Article 46 lists, including contractual clauses between the exporter and the recipient in the third country.
- A derogation for specific situations under Article 49, which is not a basis for routine production traffic.
When an inference provider processes personal data on your behalf as a processor, Article 28 requires a processing contract. Confirm the roles, purposes, subprocessors, support access and transfer arrangements. The data processing agreement (DPA) is part of that evidence; its existence alone does not prove compliance.
Per-token APIs with European processing
Three providers sell open models per token with processing in Europe and publish the details in their own documentation. Before the roster, the load-bearing caveat: on a multi-model platform, region is a property of each model, not of the provider. A provider-level European claim is checked model by model, against the record for the specific model string you call. That applies to every entry below, ours included. Our earlier breakdown of which open-weight models are actually hosted in Europe walks through the verification steps.
| Provider | API | OpenAI-compatible | Where processing happens | DPA |
|---|---|---|---|---|
| Mistral AI | La Plateforme regional API | Yes, chat completions | EU endpoint covers EU and EFTA countries; control plane remains global | Check the DPA and subprocessors for the selected service |
| Scaleway | Generative APIs | Yes, OpenAI client libraries | Paris for the published serverless offer | Request the applicable DPA and retention terms |
| Nebius | Token Factory | Yes, OpenAI-compatible API | EU or US placement depends on deployment; confirm the endpoint | Request the applicable DPA and retention terms |
Mistral’s api.eu.mistral.ai endpoint covers EU and EFTA countries, which is broader than a strict EU-only requirement. Its control plane remains global. Regional inference costs 1.1 times standard list prices. Only models available in that region are served, and Agents, Batch and Files are unavailable. Check the regional model list before migrating.
Provider pages checked on 1 October 2026 show Scaleway’s Paris GLM-5.2 rate at €1.80 input and €5.50 output per million tokens. Nebius describes EU or US placement for Token Factory deployments. Confirm the chosen region and retention terms in the endpoint configuration and contract, rather than inferring them from the company’s address.
Disclosure: we publish this article and sell per-token inference in Europe, so read every row above, and every claim below, with that in mind. The market is small and interconnected. The point of this piece is to make each provider's region claim checkable, including ours.
GPU platforms you deploy onto yourself
Renting GPU virtual machines in an EU region gives you control over the serving stack and storage configuration. It does not guarantee residency by itself: logs, backups, external tools, telemetry and support access can cross borders. You also own capacity planning, updates and incident response. Compare that operating cost with the managed API’s token bill.
- Serving stack: engine selection, quantisation, KV-cache sizing, and the CUDA OOM debugging that comes with tight VRAM.
- Capacity planning: peak concurrency versus GPU count, since dedicated capacity is billed based on the instance type and usage duration whether or not it receives traffic, as Scaleway's dedicated deployment billing shows.
- Operations: health checks, model rollouts, driver and CUDA version pinning, and security patching.
For a team without a platform engineer, that overhead is the product you were buying from Together AI in the first place. If you are weighing self-hosting against a managed API for latency reasons as well, our comparison of Groq alternatives in Europe covers the latency side of that trade-off for European workloads.
US providers offering an EU region
Together’s dedicated offering includes regional deployment choices. Verify the exact serverless or dedicated product separately, including its model, region, overflow routing and contract. A region offered for dedicated capacity should not be assumed to apply to shared per-token inference.
- It changes where the compute sits, and with it the latency for European users.
Whether a US provider's EU region is sufficient for your processing is a legal question that deserves its own analysis. The engineering point for this roster is narrower: if the requirement handed to you says processing in Europe, verify it for the specific product tier and the specific model, not for the provider's region list. The next two questions make that verification concrete.
Two questions that make a claim checkable
Every region claim in this article reduces to two questions you can put to any provider, in writing, and attach to a procurement file. They are deliberately narrow, because vague sovereignty language is where vendor claims and engineering reality drift apart.
- Where does inference execute for the specific model I need? Not the platform, not the company: the model string. Ask for the region on that model's own record, and for the hostname your requests hit.
- Does it fail over anywhere else under load? A regional endpoint that silently routes overflow to another geography does not keep the promise your DPA describes. Mistral's own documentation tells you to log the endpoint hostname, model ID and request identifier for each regional request so the processing location stays auditable.
Lyceum's answer to both questions is per model. Region is stated on each model's own record rather than as a platform-wide claim, and prompts and outputs are processed but not stored and never used for training, a zero data retention posture the company self-asserts. A DPA is available on request. We publish no SLA for the serverless tier; availability is tracked publicly at status.lyceum.technology. If you are pressure-testing any provider's sovereign claim, including ours, the checklists in how to vet a sovereign EU claim and verifying zero data retention give you the questions in procurement-ready form.
The catalogue trade you are making
Compare the model strings you actually call, rather than catalogue totals. A replacement can cover your production needs even if its overall catalogue differs. Record the required chat, coding and embedding models, then evaluate substitutes for anything missing.
- List the model strings your production and staging code calls today.
- For each, check the replacement provider's record for that exact string: region, context window, cached-input rate, price per 1M tokens.
- Check the fallback: if a model is missing, what is the nearest substitute, and has anyone on your team evaluated it?
- Re-run your eval suite on the substitutes before the switch, not after.
Choosing by the model you actually need
Lyceum Serverless Inference offers pre-hosted models on a shared endpoint, billed per token. For supported chat-completions calls, configure the base URL, API key and model identifier in the OpenAI client. Then test tool calls, streaming, reasoning fields, context limits and usage accounting before moving production traffic.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["LYCEUM_API_KEY"],
base_url="https://api.lyceum.technology/openai/v1",
timeout=30.0,
max_retries=0,
)
response = client.chat.completions.create(
model="z-ai/glm-5.2",
messages=[{"role": "user", "content": "Hello"}],
max_tokens=2048,
)
print(response.choices[0].message.content)| Model (API string) | Region | Input per 1M | Cached input per 1M | Output per 1M |
|---|---|---|---|---|
| GLM-5.2 (z-ai/glm-5.2) | EU, listed 1 October 2026 | $1.50 | $0.38 | $4.50 |
| DeepSeek-V4-Flash (deepseek/deepseek-v4-flash-0731) | EU, listed 1 October 2026 | $0.25 | $0.06 | $0.30 |
| Kimi-K3 (moonshotai/kimi-k3) | EU, listed 1 October 2026 | $3.00 | $0.75 | $15.00 |
| MiniMax-M3 (minimax/minimax-m3) | EU, listed 1 October 2026 | $0.40 | $0.10 | $2.00 |
These prices and EU hosting labels were checked in the Lyceum dashboard on 1 October 2026. Prices are USD per million tokens. Cached-input pricing applies only to cache hits, which are best effort. Recheck the exact model record and current /models response before use.
The documented limits, stated plainly: Serverless Inference is self-serve and carries no SLA, no availability tier and no service credit. No uptime figure is published for it; availability is tracked at status.lyceum.technology, and contracted options are a sales conversation.
Check the model, processing region and current rate, then run a staging comparison before changing production routing. The Lyceum dashboard lists the current model catalogue.