AI This article was created with the help of AI.
Sovereignty is four requirements, not one
When engineering teams decide to self-host open-source large language models, the justification given to leadership is almost always data sovereignty. In practice, bundling every infrastructure decision under the banner of sovereignty obscures the actual technical and legal trade-offs. What teams call sovereignty is actually four distinct operational requirements, each with its own constraints, risks, and implementation paths.
To evaluate whether your organization needs to run its own inference hardware, you have to separate these four dimensions explicitly: processing location, operational control, absence of a third party in the data path, and portability.
- Processing location: the verified regions and paths used for inference and related data.
- Operational control: which software, configuration and hardware decisions your team controls.
- External provider access: which entities can process or access payloads, including cloud hosts and support providers.
- Portability: whether license, weights and runtime compatibility let you migrate the workload.
Most infrastructure leads assume that satisfying any one of these constraints requires satisfying all four through complete in-house hosting. That assumption is flawed. Conflating processing location with stack operation leads engineering teams to take on massive operational burdens that provide no additional compliance or architectural value beyond what modern European cloud infrastructure already delivers.
What a managed European service satisfies
A managed service may meet location and portability requirements, but neither follows automatically from European hosting. Check the particular model, deployment, license and contractual controls.
An EU region can satisfy a processing-location requirement only for the specific deployment and data paths covered by the service. Verify inference routing, storage, logs, backups, support access and subprocessors against your contract and DPA. A European company address alone is insufficient.
Portability is addressed when a managed service relies on open-weight architectures rather than closed, proprietary models. While true open source implies full freedom to modify and share per Version 1.0 of the Open Source Initiative's Open Source AI Definition, many modern models are open weights with specific usage conditions. You must check the specific license of the exact model version you intend to use rather than assuming a whole family is identically licensed.
| Sovereignty Requirement | Managed EU Inference | Self-Hosted Stack |
|---|---|---|
| Processing Location | Depends on verified selected route and contract | Depends on actual hosting |
| Portability | Conditional on license and runtime | Conditional on license and runtime |
| Operational Control | Managed scope | Internal responsibility |
| Third-Party Absence | External providers involved | Owned hardware needs separate verification (rented VPCs involve external providers) |
Migration to a private cluster or another regional vendor depends on more than just having the weights. Your ability to move the workload relies on the model license, having access to the exact revision, compatible hardware, matching tokenizers or chat templates, and API feature compatibility between providers. You avoid proprietary dependencies, but achieving portability requires careful alignment of these components.
The requirement no managed option can meet
A strict organization-specific requirement to exclude every external provider needs separate assessment. Geographic location alone does not establish this boundary.
A processor processes personal data on behalf of a controller under Article 4(8). Such a processor is excluded from the Article 4(10) definition of a third party. Article 28 sets conditions for using processors; it does not generally prohibit them. Determine the provider role from the actual processing relationship.
Record the roles, purposes, retention, access and security controls in the relevant agreement. Some organization-specific policies or contracts exclude external providers; assess the actual requirement rather than assuming GDPR creates a general ban.
If your policy excludes every external provider from the data path, assess a verified deployment on hardware you own and control. A rented cloud VPC still involves the cloud operator and potentially support providers. Offline operation, bespoke audit controls, custom models and economics can also justify self-hosting.
What operational control costs in people
The fourth requirement is operational control. Teams often choose to self-host because they want total authority over every runtime setting, compiler flag, and hardware allocation. However, operational control is not a free architectural feature; it is an ongoing tax on engineering resources.
Running production-grade inference requires deploying high-throughput serving engines like vLLM. The SOSP 2023 paper that introduced PagedAttention reports that vLLM improves the throughput of popular LLMs by two to four times at the same level of latency compared to the state-of-the-art systems of the time, such as FasterTransformer and Orca, by achieving near-zero waste in KV cache memory. Realizing those performance gains in a live production environment requires continuous tuning of configuration parameters.
When self-hosting, your platform engineers must manually configure and maintain the engine arguments that control the behavior of the vLLM engine, which for online serving are passed to vllm serve: the documented set runs from the data type for model weights and activations (--dtype) and the random seed for reproducibility (--seed) through to memory and parallelism settings. These include tensor parallel sizes, pipeline parallel execution schedules, GPU memory utilization limits, max model sequence lengths, swap space allocations, and specialized attention backends. Miscalculating any of these parameters can cause sudden CUDA Out-of-Memory (OOM) crashes during traffic spikes.
- Cluster node health: Monitoring GPU health across nodes, resolving InfiniBand interconnect degradation, and recovering from silent hardware failures.
- Runtime updates: Managing rolling upgrades of CUDA toolkits, NVIDIA drivers, and engine binaries without dropping live websocket connections or corrupting KV caches.
- Autoscaling and batching: Building dynamic continuous batching schedulers and handling cold-start latency when scaling GPU nodes from zero under variable load.
- Observability and debugging: Instrumenting compiler-level profiling to diagnose slow generation loops, engine lockups, and KV cache evictions.
Staffing depends on availability targets, traffic scale, infrastructure complexity and your team's existing skills. Include on-call work, maintenance and incident response when comparing managed service fees with self-hosting costs.
Managing version changes
Version control is another reason to operate your own serving deployment. Recording exact model revisions, container digests and generation settings reduces unexpected changes, but does not guarantee identical outputs.
Managed multi-tenant inference catalogues frequently update their backend serving software, quantization routines, and default model checkpoints to optimize global server utilization. A minor patch in a serving framework or an unannounced update to a model tokenizer can subtly shift token distribution probabilities, changing the output for identical prompts. In heavily regulated industries where model behavior must be audited and verified, these changes introduce substantial compliance risk.
Pin model revisions, serving versions and container digests where supported. Hardware, kernel execution, batching and numerical precision can still affect results. Plan security updates and hardware lifecycle changes, then rerun your evaluations.
| Stack Layer | Managed Shared API | Self-Hosted Deployment |
|---|---|---|
| Model Checkpoint | Subject to provider catalogue deprecations | Pinned to exact Git commit and weight hash |
| Inference Engine | Upgraded automatically by provider | Locked to a specific tagged release |
| Engine Arguments | Standardized across multi-tenant cluster | Customized per workload (e.g. batch size, max tokens) |
| Hardware Architecture | Allocated dynamically by platform | Subject to underlying hardware and kernel variances |
For audited workloads, record the full tested configuration and change-control process. A frozen checkpoint alone cannot guarantee zero behavioral drift across a multi-year period.
Where a dedicated endpoint sits between them
Between the operational overhead of managing raw GPU clusters and the rigidity of shared multi-tenant APIs lies a pragmatic middle ground: dedicated inference endpoints. Dedicated endpoints reserve resources and shift some operations to a provider; the exact control and support boundary depends on the offering.
A dedicated endpoint reserves compute for a workload. Lyceum documents choices for model identifier, hardware and scaling; confirm supported checkpoint selection, engine settings and network access before relying on them.
Dedicated compute can reduce resource contention, but reserved GPUs do not automatically provide arbitrary engine pinning, private networking or zero latency jitter. Verify the required controls and lifecycle commitments with the provider.
By adopting dedicated endpoints, you avoid the heavy burden of kernel-level debugging, hardware fault recovery, and cluster autoscaling. For teams exploring dedicated versus shared GPU inference, this structure provides the ideal balance between operational control and infrastructure efficiency.
Making the independence decision
Deciding between self-hosting and managed EU inference requires an honest assessment of which sovereignty constraints your organization actually has, rather than adopting self-hosting as a default reaction to regulatory pressure.
Verify model-specific processing paths and contractual controls for residency. Assess license/runtime compatibility for migration. Consider owned infrastructure for strict provider exclusion, and weigh offline access, audit, custom models and economics. Dedicated endpoints may fit if the offered model, hardware, scaling and access controls meet your requirements; verify required pinning, isolation and service terms before choosing.