The naive comparison always favours self-hosting
A GPU hourly rate divided by peak throughput can make self-hosting look cheap. A useful comparison also includes idle hours, supporting infrastructure and engineering. Whether self-hosting wins depends on the workload and the complete cost model.
Use measured throughput at your latency target. A benchmark at constant load is different from a production trace containing quiet periods, retries and bursts. Label the measurement window so the spreadsheet treats idle time consistently.
If active throughput is q tokens per second and the service is active for fraction U of billed time, average throughput is q × U. If q already measures the complete billing window, do not multiply by U again. A 1/U cost multiplier applies only when active throughput and fixed hourly cost are held constant.
- GPU hours you pay for but fill with no traffic, nights and weekends included
- Storage and network egress on your own account, not the provider's
- Engineering time to select an engine, quantise the model and ship a deployment
- On-call rotation and upgrade load for the serving stack, permanently
- Autoscaling you now operate yourself, with controllers you install and maintain
State plainly: there is no general crossover figure. It depends on your model, your context length, your traffic shape and the utilisation you actually achieve, and you compute your own in the last section of this guide. How model size and device count bend that curve is covered separately in the token break-even by model size. One disclosure before the arithmetic: Lyceum publishes this article and sells Dedicated Inference, which competes with both extremes. The method below is the same either way.
What per-token pricing actually buys you
Read the per-token price as a capacity-management fee, not a markup. When you pay per token, someone else owns the gap between peak and average traffic. The GPU that serves your 9 a.m. spike is provisioned at 3 a.m. too, and the provider carries it. You are buying tokens, but you are also buying the scheduling, batching and memory management that keep expensive hardware productive across wildly uneven load.
That engine work is the part the naive comparison treats as free. Sustained LLM throughput is set by batching and by how well the serving system manages the key-value cache, which grows and shrinks with every request and wastes memory through fragmentation when handled poorly. PagedAttention, the mechanism behind vLLM, eliminated that waste and improved throughput of popular LLMs by 2 to 4 times at the same latency compared with the then state of the art. That improvement is real, but it had to be engineered. A per-token buyer inherits it; a self-hosting team has to build and keep it.
- Idle GPU time between your requests, absorbed by the provider's multi-tenant pool
- Continuous batching and KV-cache memory management, the levers that set sustained throughput
- Prompt caching, where repeated context is billed at the cached rate instead of recomputed
- Engine upgrades and model version bumps as the open-source stack moves
- Capacity headroom for traffic bursts you did not forecast
The per-token side of the crossover also moves with your own traffic shape. Traffic that can wait for a scheduled window does not need a real-time endpoint, and async batch inference runs at half the list price. The trade-off between the two modes is laid out in batch versus real-time inference pricing.
Costing the self-hosted side in full
The self-hosted column in most comparisons contains one number: the GPU rate. The full cost of running an open model yourself has at least four more lines, and three of them recur every month. Frame the decision as total cost of compute, compute plus storage plus engineering plus provisioning waste, not an hourly GPU rate.
| Cost item | What it includes | Why the naive sheet misses it |
|---|---|---|
| GPU time including idle | A reserved GPU bills around the clock, at 10 percent load and at 90, nights and weekends included | The sheet divides by peak throughput, which implies the GPU never idles |
| Storage and network | Model weights, KV-cache offload, checkpoints, and egress on your own account | Bundled into the per-token price on the API side, invisible as a line item |
| Engineering time to build the stack | Engine selection, quantisation, deployment, load testing, CI for model updates | One-off on the spreadsheet, recurring in reality |
| On-call and upgrades | Pager duty for OOM crashes and node failures, engine and driver upgrades, security patches | Exists on neither side of the naive comparison |
| Autoscaling you operate | Controllers to install, configure and keep alive so replicas follow demand | Assumed to be a solved problem; it is a small permanent project |
The last line deserves its own arithmetic. When you run inference yourself, scaling to demand means running a HorizontalPodAutoscaler or VerticalPodAutoscaler against your deployment, and the VPA in particular is an add-on you or a cluster administrator must deploy before you can use it, with its own prerequisites to maintain. Every replica you keep warm to absorb a burst is a GPU you pay for whether a token moves through it or not.
Engineering time is the line procurement teams underweight most. Building a serving stack that holds a latency target under real concurrency is weeks of senior engineering, and the stack does not stay built: engines release, models change, drivers and CUDA versions move. Amortise that time over a realistic lifetime and it lands on the cost side of the crossover like any other fixed cost.
Utilisation decides your effective token cost
A reserved GPU costs the same at ten percent load as at ninety. That single fact makes utilisation the deciding variable in the whole comparison. Effective cost per token is the rented cost divided by the tokens actually served, so the same hardware at the same rate produces wildly different unit economics depending on how much of the day it has work.
There is no single enterprise utilisation percentage that belongs in your budget. Measure the paid capacity and served workload in your own system. GPU hardware utilisation, busy-time fraction and achieved token throughput are different metrics.
| Illustrative active-time fraction | Fixed GPU cost per token versus full activity |
|---|---|
| 100% | 1× |
| 50% | 2× |
| 20% | 5× |
| 5% | 20× |
These rows illustrate 1/U with unchanged active throughput and hourly cost; they are not measured industry averages. Real throughput can also change with batching, context, concurrency and the latency target. Recalculate using your traffic.
Measuring sustained throughput yourself
Throughput is measured, not borrowed. A tokens-per-second figure quoted without model, context length, batch size and concurrency describes someone else's workload, and pasting it into your crossover produces a confident wrong answer. Measure sustained throughput at your own context length, your own concurrency and your own latency target, and use consistent definitions when you do, because benchmarking tools disagree on them. NVIDIA's NIM benchmarking documentation defines the three that matter: time to first token (TTFT), which includes queuing, prefill and network latency; inter-token latency (ITL), the average time between consecutive output tokens; and total tokens per second per system, which rises with concurrent requests until GPU compute saturates and can then decrease.
- Fix the model, precision, runtime and representative input/output lengths
- Replay a representative workload, including the quiet periods you pay for
- Record successful tokens, errors, total billed time and latency percentiles
- Compute successful tokens divided by the complete billing window for wall-clock throughput
- Use that result directly for capacity costing. Apply an activity fraction only if throughput was measured exclusively during active service
For hardware-level context, MLPerf Inference: Datacenter publishes comparable results across platforms, with each row carrying its submitter, software stack, accelerator and count, so use it as a hardware reference cited with the round and configuration, never as a throughput promise for your workload. For sizing, the H100 SXM specification gives 80GB of HBM3 memory at 3.35TB/s of memory bandwidth, which bounds how large a model and how much KV cache a single card holds. The specification is a specification; it is not a price and not a throughput figure.
The dedicated endpoint in the middle
A managed dedicated endpoint reserves capacity while the provider operates the serving stack. Lyceum’s documented workflow uses a Hugging Face model identifier, supported hardware and replica settings. Confirm model support and API features before assuming that an existing application will work unchanged.
- A supported Hugging Face model and any required model-access token
- Hardware sized for weights, cache and expected concurrency
- Replica bounds and the required isolation and location terms
- Scaling behaviour, billed idle time and cold-start latency
- Integration tests for the API features your application uses
Lyceum documents scale-to-zero after one hour of inactivity, followed by a cold start when traffic returns. Include that paid idle window and startup behaviour in the comparison. Dedicated service commitments depend on the signed contract; serverless inference has no contractual service-level agreement.
Finding your crossover in tokens per month
Choose a consistent unit such as one million tokens at a fixed input/output/cache mix. Let F be monthly fixed self-hosted cost, p the managed API cost per unit, and v the self-hosted variable cost per unit. The crossover is F / (p − v) units per month, provided p > v. Do not add electricity separately if it is already included in rented GPU hours.
- Calculate API cost from the same fresh-input, cached-input and output mix used in the self-hosted workload
- Include GPU rental, storage, network, amortised setup and recurring operations in the appropriate fixed or variable bucket, once
- For an illustrative F of $3,000, p of $2 per million mixed tokens and v of $0.50, crossover is 2,000 million tokens per month
- If p is less than or equal to v, this model has no positive fixed-cost crossover
- Check whether the rented capacity can serve the crossover volume within the latency target. Add hardware costs when capacity must grow
- Reprice a dedicated endpoint with its actual replica billing, idle delay and cold starts
Two assumptions move the result more than everything else: the utilisation you will actually sustain, and the engineering cost you assign to the stack. Sensitivity-test both before taking the number to a budget review. And note where the arithmetic is not the decision. Data handling requirements can rule a US-hosted API out regardless of price. Latency control can rule shared endpoints out for interactive workloads. Model availability can rule self-hosting out when the weights are not open. These overrides cut in both directions, and they should be written down next to the crossover, not discovered after the migration.
The decision is also reversible, which lowers the cost of testing it. Capacity can be added or removed with 2 to 3 weeks notice and there is no penalty for scaling down, and a proof of concept typically runs 4 to 8 weeks on real workloads, so you can measure your crossover on production traffic before committing. Measure your sustained throughput and utilisation, then price all three options rather than two, starting with the dedicated endpoint.