AI This article was created with the help of AI.

The crossover is not one number

Engineering teams evaluating open-source language models frequently seek a single break-even volume for self-hosting. The standard inquiry asks whether generating ten million or fifty million tokens per month makes renting dedicated graphics processing units cheaper than consuming a managed inference endpoint. Framing infrastructure economics around a universal crossover volume is a fundamental misconception. The token break-even point is not a fixed threshold across open architectures; it is a dynamic curve dictated primarily by model parameter size and the physical hardware required to host it.

The Open Source AI Definition published by the Open Source Initiative sets out the freedoms to use, study, modify and share an AI system, and asks that the preferred form for making modifications be available. When you deploy open weights directly, you assume the operational and capital constraints of hosting those parameters in physical silicon. Depending on model architecture, the point where self-hosting becomes cheaper than managed token consumption shifts by orders of magnitude.

To establish a predictable framework for infrastructure planning, we evaluate the self-hosting crossover across three distinct architectural classes:

  • Small dense models (7B to 9B parameters) operating on a single discrete accelerator.
  • Large dense models (70B parameters) requiring tensor parallelism across multi-GPU nodes.
  • Mixture-of-Experts (MoE) architectures where total memory footprint diverges drastically from per-token compute execution.

Because provider token pricing scales alongside compute requirements, both sides of the economic equation move simultaneously. Understanding how device counts dictate fixed hosting baselines is essential before committing to self-managed infrastructure.

Why device count drives the fixed cost

The primary economic driver in self-hosted inference is the step-function pricing curve of hardware provisioning. Unlike cloud software that scales continuously with CPU and memory allocations, Large Language Model (LLM) serving requires full model weights and dynamic attention state to reside entirely within high-bandwidth accelerator memory. A model cannot be partially provisioned; it demands a discrete integer count of accelerators before handling its first request.

Serving frameworks such as vLLM use PagedAttention to manage the key-value cache in pages; the authors report near-zero waste in KV cache memory and a throughput improvement of two to four times over prior systems at the same level of latency. The project's own repository documents that same paged memory management alongside continuous batching and tensor parallelism as core serving features, and its engine arguments expose runtime settings such as max_model_len and gpu_memory_utilization so engineers can tune memory headroom and batching efficiency. Yet software optimization cannot bypass physical memory capacity. When a model's memory footprint exceeds a single card and requires tensor parallelism, your baseline operational expenditure instantly doubles or quadruples.

When evaluating transformer GPU memory requirements, the baseline fixed cost represents an inescapable floor. If a model requires four 80GB SXM accelerators to run comfortably with adequate context caching, you pay the continuous hourly rate for four devices regardless of whether your incoming query volume utilizes 5 percent or 95 percent of that compute capacity.

A small model on a single device

Small dense architectures, spanning from 7B to 9B parameters, represent the most economically accessible entry point for self-hosting. At two bytes per parameter, an 8B model held in FP16 needs about 16GB of memory for its static weights alone, before any cache is allocated. In FP8 that checkpoint footprint falls to about 8GB, and at INT4 to about 4GB, with lower precision reducing memory at some cost in accuracy.

Because an 8B model fits comfortably onto a single accelerator such as an NVIDIA L40S (48GB) or NVIDIA A100 (80GB), the hosting configuration requires no inter-GPU communication or tensor parallel overhead. The remaining memory headroom allows extensive concurrency when calculating KV cache memory for high-throughput batching.

Model ClassStatic weight memory (16-bit)Devices needed at 80GB per acceleratorCrossover driver, quantified
Small dense (8B)16GB for an 8B model in FP16One, since 16GB of weights sits well inside a single 80GB cardOne device-hour against the per-token rate, with the rest of the card's memory available for KV cache and batching
Large dense (70B)140GB in FP16, falling to 70GB in FP8 and 35GB at INT4Two 80GB cards for FP16 weights alone, with more added for context headroomSeveral device-hours against the per-token rate, so a multiple of the 8B fixed cost before a single request is served
Large MoE (Mixtral 8x7B, about 45B total parameters)More than 90GB in float16, set by total rather than active parametersTwo 80GB accelerators to hold the FP16 model, or one at 8-bit (more than 45GB)A 90GB memory floor against per-token rates priced off 12B-dense decode speed, since two experts run per token

On a single accelerator, continuous saturation translates directly into cost advantages. Because the fixed hourly cost is restricted to a single device, even moderate enterprise workloads generating consistent daily traffic can cross the threshold where dedicated infrastructure costs less than pay-as-you-go token pricing. For continuous, high-concurrency internal pipelines, small models clear the self-hosting break-even quickly.

A large dense model across several

Scaling up to a 70B parameter dense architecture alters the break-even math. A 70B model in standard 16-bit precision requires approximately 140GB of memory merely to load static weights, which is more than any single 80GB accelerator holds, so the deployment starts at two devices. Those figures cover the checkpoint only, not the runtime memory reserved for kernels or CUDA graphs or the growing key-value cache, so production deployments with long context windows commonly add further devices to avoid out-of-memory faults. Quantization pulls the floor down, to about 70GB in FP8 and 35GB at INT4, at some cost in accuracy.

Serving a 70B model in FP8 halves the weight footprint and lets modern high-end accelerators sustain high generation rates under saturated batch loads. However, the fixed hourly cost has still scaled with the number of devices you had to provision relative to a single-card setup. While commercial API providers charge higher per-token rates for 70B models, the steep initial hosting baseline pushes the crossover threshold significantly higher.

When choosing between dedicated versus shared GPU inference, large dense models require sustained utilization to justify their physical infrastructure. If token demand arrives in unpredictable bursts rather than continuous streams, the cumulative cost of running multi-GPU nodes through idle periods rapidly outweighs the per-token expense of managed endpoints.

Where mixture-of-experts reverses the answer

Mixture-of-Experts (MoE) architectures completely invert conventional self-hosting calculations. In an MoE model, the total parameter count determines the required memory capacity, but only a small fraction of those parameters (the active parameters) execute during any individual token generation step. For example, a model might possess over 200 billion total parameters while activating only 20 billion to 30 billion parameters per forward pass.

For an API provider, MoE models are exceptionally cost-efficient to serve. Because token pricing correlates with floating-point operations (FLOPs) and compute execution time, managed providers can price MoE queries near the level of small dense models. Conversely, self-hosting an MoE model requires keeping all hundred-plus gigabytes of expert weights resident in GPU memory simultaneously across an entire cluster of four to eight high-bandwidth accelerators.

This structural dichotomy creates an economic paradox: MoE models offer some of the cheapest per-token rates on managed APIs, yet represent the most capital-intensive architectures to self-host. A team attempting to self-host a large MoE model must generate massive, non-stop token throughput to offset the multi-device hardware floor against low API rates.

Low utilisation costs most at the top end

Evaluating break-even models under the assumption of 100 percent sustained hardware saturation is a critical planning error. In real-world enterprise environments, inference traffic follows diurnal business cycles, developer working hours, and intermittent batch runs. When compute capacity sits idle, the effective cost per generated token rises sharply.

MLPerf Inference: Datacenter measures each system with a standard load generator that issues requests in defined scenarios, with latency constraints attached to the throughput a submitter may report. In production, maintaining strict tail-latency targets prevents running GPUs at theoretical maximum saturation. Well below full average capacity utilization, a multi-device cluster hosting a 70B or MoE model spreads its steep hourly expense over relatively few tokens.

Utilization LevelEffective Cost MultiplierSingle-Device ImpactMulti-Device Cluster Impact
100% Saturation1.0x (Baseline)Optimal economic crossoverHigh token efficiency
50% Average Load2.0x Effective CostModest cost penaltySubstantial capital waste
20% Off-Peak Load5.0x Effective CostManageable dollar impactSevere economic penalty
5% Idle Baseline20.0x Effective CostLow idle baseline spendMaximum unrecovered overhead

The financial penalty of underutilization scales directly with device count. An underutilized single accelerator represents a manageable monthly variance, whereas an underutilized eight-GPU cluster constitutes a major operational liability that invalidates self-hosting assumptions.

Placing your model on the curve

Determining whether to self-host or rely on an API requires plotting your specific architecture, parameter count, and expected concurrency against the hardware sizing curve. The decision is never binary; it is an architectural calculation that balances the fixed hourly cost of required silicon against live per-token pricing for managed endpoints.

  1. Map the exact model memory footprint, including static weights in your target precision and dynamic KV cache headroom, starting from the model card that documents the model's architecture, intended uses and limitations.
  2. Identify the minimum integer device count required to host the workload without risking out-of-memory failures.
  3. Calculate your expected average capacity utilization across 24-hour cycles rather than assuming constant peak throughput.
  4. Compare the fixed monthly cost of dedicated devices against the projected token expenditure on managed endpoints.

For enterprise teams deploying large open models that require isolated compute, data sovereignty, and custom runtime optimizations without maintaining physical clusters, Lyceum Technology provides Dedicated Inference. By deploying isolated private endpoints on dedicated European hardware, you eliminate noisy-neighbor interference and secure predictable capacity tailored directly to your production workload.

Place each model on the curve separately, then size the endpoint for the ones that stay managed with dedicated inference endpoints.