Where V4 Flash and V4 Pro can run today
This page explains where Lyceum lists DeepSeek V4 Flash and V4 Pro as EU-hosted, what each costs and how to try the API. DeepSeek’s first-party privacy policy describes processing and storage in China. A separately hosted model has a different data flow, which your team should assess before sending personal data.
Both DeepSeek V4 Flash and DeepSeek V4 Pro are served on serverless inference infrastructure, hosted in the eu-north1 region inside the European Union. The dashboard model catalogue, read on 30 September 2026, lists both text models in its EU-hosted section. You call them through one OpenAI-compatible endpoint, and the model string selects the model: deepseek/deepseek-v4-flash-0731 for Flash and deepseek/deepseek-v4-pro-0813 for Pro. Each carries a 1M-token context window and accepts text input, with reasoning on by default and switchable per request.
| Model | API model string | Hosting region | Context window |
|---|---|---|---|
| DeepSeek V4 Flash | deepseek/deepseek-v4-flash-0731 | eu-north1 (EU-hosted) | 1M tokens |
| DeepSeek V4 Pro | deepseek/deepseek-v4-pro-0813 | eu-north1 (EU-hosted) | 1M tokens |
The company catalogue records eu-north1 as the EU-hosted tag for these 2 model IDs. It does not identify a particular city or data centre. Record the model ID, region label and check date, then confirm processing locations, remote access and onward transfers in the data processing agreement. The dashboard label alone does not establish your compliance.
Per-token pricing, model by model
Both models are billed per token, with no minimum spend, and prices are published per model excluding VAT. The catalogue lists DeepSeek V4 Flash at $0.25 per 1M input tokens, $0.06 per 1M cached input tokens and $0.30 per 1M output tokens, and DeepSeek V4 Pro 0813 at $2.00, $0.50 and $4.00 respectively.
| Price component | DeepSeek V4 Flash | DeepSeek V4 Pro |
|---|---|---|
| Input, per 1M tokens | $0.25 | $2.00 |
| Cached input, per 1M tokens | $0.06 | $0.50 |
| Output, per 1M tokens | $0.30 | $4.00 |
Flash costs about one eighth of Pro on input tokens and one thirteenth on output tokens. Both also publish a lower cached-input rate. Prompt caching is automatic and best-effort: keep stable prefixes together, but budget for cache misses and read the usage fields defensively. The saving depends on how much input actually reaches the cache.
Reasoning can consume the output allowance as well as the final answer, so budget for both. Flash has lower input, cached-input and output rates than Pro, which makes it cheaper for the same token counts in every mix. Total cost per accepted result can still differ because models generate different amounts of output, need different retries and achieve different success rates. Measure those differences on your workload.
Choosing Flash or Pro by request type
DeepSeek reports V4 Pro as a 1.6T-parameter mixture-of-experts model with 49B active parameters per forward pass, and V4 Flash as 284B total with 13B active. Both expose a 1M-token context. These architecture figures do not establish task quality, latency or throughput on Lyceum. Compare accepted results, retries and cost on your own requests. For the smaller model’s architecture and vendor benchmarks, see DeepSeek V4 Flash specs and how to run it.
- Consider Pro for difficult reasoning or coding tasks where its measured success rate justifies the higher rate
- Consider Flash for high-volume extraction, classification or short tool turns where it meets your quality threshold
- Test both on representative traffic and compare cost per accepted outcome, including failures and retries
If your traffic contains both kinds of requests, splitting them automatically is its own engineering problem, and the dedicated guide on routing requests by difficulty covers it. This article stays with the static choice: pick one string per deployment, measure, and revisit.
Where the first-party DeepSeek service stores data
DeepSeek’s privacy policy names Hangzhou DeepSeek Artificial Intelligence Co., Ltd. as the controller, with a registered address in China. The policy describes the data handling of its first-party services; assess that service separately from another provider hosting the open weights.
- The policy states that to provide its services, DeepSeek directly collects, processes and stores personal data in the People's Republic of China.
- The personal data it collects includes your prompts and inputs: text input, uploaded files, chat history and other content you provide to the model, alongside device and network data such as IP addresses.
- The policy states that personal data may be used to train and improve DeepSeek's technology and machine-learning models.
- The policy also notes that it does not cover the processing of end-user personal data in downstream applications built on DeepSeek's open platform; the developer operating that application is the controller for that processing.
For a European enterprise, that combination is what halts the review: prompts containing customer or employee data would be a transfer of personal data to a third country, and China does not appear on the European Commission's list of countries with an adequacy decision. The open weights change nothing about the first-party service; they change what you can do instead, which is run the same model on infrastructure you can place inside the EU.
Why residency is a per-model question
The most common evaluation mistake is to assume hosting location at the provider level: a European-registered vendor, a European invoice, an EU contract. None of that determines where the GPU executes your prompt. A platform can route a request to hardware outside the EU while every commercial document stays European, and the residency obligation attaches to the payload, not the paperwork.
Both model IDs carry an EU-hosted label in the current dashboard and an eu-north1 tag in the company catalogue. Record this advertised inference location, then check contractual locations and onward processing. The broader method is described in which open-weight models are actually hosted in Europe.
- Check the model's own record or the dashboard catalogue for a region tag, per model string, before you put the string into production.
- Treat the /models endpoint as the source of truth for what currently exists; a model that is not listed is not an option, whatever an older article says.
- Do not assume a model was always hosted where it is hosted today; migration dates are generally not published, so verify at the time you deploy.
The discipline costs minutes and buys a defensible answer for your data protection officer: the exact string, the exact region tag, the date you checked it.
What running in Europe does not settle for you
The dashboard labels both model IDs as EU-hosted. That answers the advertised inference location, but it does not prove that every log, support system or sub-processor remains in the European Economic Area. For personal data, check the complete data flow and the processor agreement. Chapter V applies when the EDPB transfer criteria are met; EU inference alone is not a compliance verdict.
What it does not do is make you, the controller, compliant. The EDPB is explicit that the conditions for international transfers come in addition to the basic processing principles: you still need a legal basis for processing, security measures, data minimisation, and a contract with the recipient where it acts as your processor.
- EU hosting identifies the advertised inference location for the model ID you checked
- You still need to assess lawful basis, data minimisation, security, retention, processor terms and any remote access or onward transfers
- AI Act duties depend on your system and role; choosing a hosting region does not resolve them
Switching is a base-URL change
The OpenAI Python client accepts a base_url override. For supported chat-completions requests, you can keep the client library and change the endpoint, API key and model ID. Test streaming, tool calls, reasoning fields and limits against your application before moving production traffic. OpenAI compatibility does not promise every feature or identical behaviour.
- Point base_url at the OpenAI-compatible endpoint for the model you are calling.
- Replace the API key with your own API key, sent as a Bearer token.
- Set the model string to deepseek/deepseek-v4-flash-0731, or deepseek/deepseek-v4-pro-0813 if that is the model your workload needs.
That closes the loop for the situation this article opened with: a model your teams want, priced per token, stopped by a data-flow question the first-party service cannot answer. DeepSeek V4 Flash on Serverless Inference is the direct answer for the high-volume side of that workload: the deepseek/deepseek-v4-flash-0731 string, EU-hosted in eu-north1, billed per token with no minimum commitment, reachable from the OpenAI client you already ship. Check the region and price on the DeepSeek V4 Flash model record, then point your base URL at it.
Lyceum states that inference prompts and outputs are processed without retention and are never used for training. Its approved caching statement is GPU memory only, per session, for minutes at most. This is a company statement, without third-party attestation; request the data processing agreement if you need contractual confirmation. For Claude Code, use the separate Anthropic-compatible route at https://api.lyceum.technology/anthropic and follow the current Claude Code documentation.
from openai import OpenAI
client = OpenAI(
base_url="https://api.lyceum.technology/openai/v1",
api_key="lk_your_api_key_here",
)
response = client.chat.completions.create(
model="deepseek/deepseek-v4-flash-0731",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)