AI This article was created with the help of AI.
Two ways of allocating the same risk
When evaluating AI infrastructure costs, engineering teams often reduce the decision between per-token APIs and per-GPU-hour instances to a simple comparison of headline rates. This framing misses the fundamental mechanical reality of inference economics. The choice between paying per million tokens and renting raw GPU capacity is not merely an arithmetic preference; it is a structural decision about how your organisation allocates hardware utilisation risk.
Inference workloads require dedicated memory bandwidth and computational units to process prompts and generate auto-regressive tokens. Under a per-token pricing model, the provider absorbs the hardware utilisation risk in full. When your application experiences a traffic lull, your compute invoice drops to zero. The provider carries the financial burden of managing idle silicon, orchestrating multi-tenant cluster capacity, and packing concurrent requests into VRAM. You treat inference strictly as a variable operational expense that scales linearly with user activity.
| Dimension | Per-Token Model | Per-GPU-Hour Model |
|---|---|---|
| Cost Structure | Variable (metered per input/output token) | Fixed (metered per clock second or hour) |
| Utilisation Risk Bearer | Inference provider | Infrastructure buyer |
| Cost of Zero-Traffic Periods | Zero expense | The full provisioned hourly rate, unchanged |
| Capacity Management | Managed dynamically by the platform engine | Engineered and scaled by internal ML teams |
| Economic Sweet Spot | Low, intermittent, or unpredictable duty cycles | Sustained, continuous, high-concurrency saturation |
Conversely, renting dedicated GPU hardware by the hour converts inference into a fixed capital commitment. Once an instance is provisioned, you pay for every elapsed clock second regardless of whether the Tensor Cores are executing matrix multiplications or sitting idle. If your engineering team deploys a dedicated instance for an internal developer assistant that is active only during European business hours (40 hours per week), you are subsidising 128 hours of unutilised hardware time every week. Under this model, cutting GPU idle time becomes an operational burden for your own platform engineers rather than the cloud provider.
This dynamic is particularly pronounced with open-weight models. Version 1.0 of the Open Source AI Definition describes an Open Source AI as a system made available under terms that grant the freedoms to use, study, modify and share it, and it requires that the preferred form for making modifications include the model parameters, such as weights, alongside data information and code. Because open models provide access to weights and parameters, teams have the freedom to run them on self-managed infrastructure or consume them via pre-hosted endpoints. Deciding where to host them begins with understanding how much risk you can afford to hold.
What per-token pricing actually buys
Purchasing inference on a per-token basis is frequently viewed as paying a retail markup over raw compute. From a pure hardware perspective, this view appears plausible: providers must cover their data center overhead, hardware depreciation, and operational margins. However, what per-token pricing actually buys is complete insulation from provisioning friction, operational maintenance, and low-utilisation penalties.
When you execute a request against a per-token endpoint, you are buying into an aggregated, highly optimised serving fabric. Managing large language models in production requires sophisticated serving stacks to mitigate key-value (KV) cache fragmentation, schedule dynamic batch arrivals, and partition model weights across multiple GPUs using tensor and pipeline parallelism. If you manage dedicated instances yourself, your team must configure and maintain these orchestration layers, debug CUDA out-of-memory errors, and engineer auto-scaling policies that inevitably suffer from provisioning latency.
- Zero idle cost: You never pay for unutilised VRAM or idle compute cycles during off-peak hours, weekends, or sporadic usage dips.
- Elastic throughput: Immediate scaling to accommodate traffic spikes without waiting for node boot times or cold-start container initialization.
- Elimination of orchestration overhead: No requirement to maintain dedicated Kubernetes clusters, GPU node pools, or complex custom autoscaling logic.
- Built-in stack optimisation: Automatic utilization of modern serving techniques like continuous batching and optimised attention kernels without in-house maintenance.
For applications with irregular duty cycles, scale-to-zero savings on GPU inference quickly outweigh any theoretical per-hour hardware discount. The premium paid on the per-token rate functions as an operational hedge: you pay a predictable fee per unit of actual output rather than maintaining fixed infrastructure that bleeds capital during quiet hours. Keeping a self-hosted deployment consistent as models, weights and serving dependencies change is continuing pipeline work, and managed endpoints abstract that away entirely.
Classifying your own traffic shape
Determining whether per-token or per-GPU-hour billing makes economic sense requires analyzing the mathematical shape of your request telemetry. Inference traffic rarely arrives as a flat, predictable stream. By categorizing your workload into one of four primary traffic archetypes, you can map your demand curve directly to its optimal commercial structure.
- Steady and predictable: Workloads that run continuous 24/7 workloads, such as asynchronous document ingestion pipelines, automated synthetic data generation, or high-volume background processing. These workloads maintain a high, constant GPU load and strongly favour dedicated hourly GPU rental.
- Diurnal business-hours: Enterprise applications, internal coding assistants, and workplace productivity agents tied to human working hours. Traffic surges during local office hours and drops near zero overnight and on weekends, so the hardware sits idle for most of the week.
- Spiky and bursty: Event-driven workflows, interactive customer-facing agents, or applications subjected to sudden marketing-driven traffic surges. Peak request rates can run many multiples of the baseline for brief intervals, making static capacity planning economically prohibitive.
- Genuinely unpredictable: Early-stage product validation, exploratory internal research, and experimental features where query volumes fluctuate erratically from day to day. These workloads belong exclusively on per-token pricing until volume stabilizes.
When teams with diurnal or bursty traffic rent dedicated instances, they face an impossible trade-off: provision for the peak and waste capital on hardware that is idle for most of the day, or provision for the average and experience severe request queuing, latency spikes, or dropped connections during peak hours. Understanding your duty cycle is the first concrete step toward eliminating this waste.
The utilisation threshold that decides it
To move past qualitative arguments, you need a deterministic framework to calculate the crossover point where dedicated hardware becomes more cost-effective than per-token consumption. This threshold is not static; it is a direct mathematical function of your sustained throughput, your model's parameter footprint, and the effective hourly cost of the underlying accelerator.
The crossover calculation evaluates the effective cost per million tokens generated on a dedicated GPU versus the market per-token rate. You can calculate your hourly token generation on a dedicated instance using this formulation:
- Hourly Token Output = Average Tokens per Second (active generation) * seconds in an hour * Sustained Utilisation Percentage
- Effective Cost per Million Tokens = (Hourly GPU Rental Cost / Hourly Token Output) * one million
- Decision Rule: If Effective Cost per Million Tokens < Provider Per-Token Price, dedicated hourly billing is cheaper. If Effective Cost per Million Tokens > Provider Per-Token Price, per-token billing wins.
The critical variable in this equation is the sustained utilisation percentage. In production deployments, real-world constraints prevent dedicated GPUs from running fully saturated around the clock. Standardized industry evaluations make this explicit: MLPerf has defined four different test scenarios for its inference benchmarks, and a given scenario is evaluated by a standard load generator generating inference requests in a particular pattern and measuring a specific metric. Arrival pattern, not raw peak throughput, is what governs achievable load, so production clusters hold spare headroom rather than run flat out and average utilisation sits well below capacity. Halving your effective utilisation doubles your effective cost per token on dedicated hardware, pushing the crossover point in favour of per-token APIs.
Why batching makes per-token competitive
Engineers often assume that an intermediary provider must always charge more than self-hosted hardware because of profit margins. This intuition fails in deep learning inference because GPU hardware efficiency scales non-linearly with batch size and concurrency. A single tenant serving modest traffic cannot utilize modern GPU architectures efficiently, whereas a multi-tenant provider operates at massive continuous concurrency.
The primary bottleneck in serving large language models is memory bandwidth during the auto-regressive generation phase. When running a batch size of 1, the GPU reads entire parameter matrices from VRAM into SRAM just to generate a single token for one user. As batch size increases, the same weight matrices are reused across dozens of concurrent sequences, dramatically increasing arithmetic intensity and overall token throughput per second.
In traditional serving systems, the key-value cache memory for each request is huge and grows and shrinks dynamically, so when it is managed inefficiently that memory is significantly wasted by fragmentation and redundant duplication, limiting the batch size. The PagedAttention algorithm published by Kwon et al. addresses this by borrowing virtual memory and paging techniques from operating systems, and the authors report that vLLM, the serving system built on top of it, achieves near-zero waste in KV cache memory and improves the throughput of popular LLMs by 2-4 times at the same level of latency compared with state-of-the-art systems such as FasterTransformer and Orca. Production deployments of that engine expose its dynamic memory and scheduling behaviour through engine arguments. Because per-token providers aggregate thousands of distinct user streams into these high-throughput continuous batching pipelines, their effective cost per token is structurally lower than a private instance processing sporadic requests.
Non-cost factors that override the calculation
While the utilization formula provides a clear financial baseline, non-financial operational realities frequently override raw arithmetic. In enterprise environments, infrastructure choices must satisfy data governance policies, tail latency guarantees, and engineering resource constraints before cost optimization takes place.
Latency consistency is the first primary constraint. On a dedicated GPU instance, your application has exclusive access to the hardware queue. There are no multi-tenant noisy neighbours, no shared rate limits, and no sudden token throttling during cluster-wide demand spikes. If your application powers a synchronous, interactive user interface requiring deterministic P99 latency, dedicated instances provide predictability that shared endpoints cannot guarantee.
| Operational Criterion | Per-Token Serverless | Dedicated Private GPU |
|---|---|---|
| Infrastructure Maintenance | Zero engineering maintenance required | Requires ongoing driver, CUDA, and cluster operations |
| Tail Latency (P99) | Subject to provider multi-tenant load balancing | Fully isolated and deterministic |
| Data Residency & Isolation | Requires verified regional tenancy guarantees | Complete hardware and network isolation |
| Customisation Depth | Standard open-weight architectures | Full freedom to run custom CUDA kernels and LoRAs |
| Cold Start Delays | Zero cold starts on pre-hosted models | Provisioning delays when scaling instances |
The second critical factor is operational overhead. Running self-managed GPU instances requires platform engineering talent to handle kernel updates, health monitoring, dynamic autoscaling, and driver crashes. For teams focused on shipping enterprise AI applications rather than building infrastructure tooling, offloading operations to dedicated inference endpoints or serverless APIs eliminates substantial platform maintenance costs.
Splitting base load from peak
Rather than treating the choice as an either-or dichotomy, mature engineering organizations adopt a hybrid architecture. The most efficient deployment model splits traffic by duty cycle, directing the predictable, continuous floor of your workload to reserved hardware while routing variable spikes to on-demand serverless capacity.
In this architecture, your infrastructure team analyzes 30-day telemetry to identify the absolute minimum baseline traffic that runs 24/7. You provision dedicated GPU capacity sized precisely to handle this steady-state volume at high sustained utilisation, which is where the hourly rate is most efficient. When traffic surges above this threshold during peak business hours, your routing layer automatically offloads the overflow to serverless endpoints.
- Step 1: Quantify your steady-state baseline in continuous tokens per second from historical Prometheus or gateway metrics.
- Step 2: Provision dedicated instances sized strictly for this floor load, keeping target utilisation high enough that the hourly rate beats the per-token rate.
- Step 3: Deploy an intelligent inference gateway to monitor local queue depths and latency metrics in real time.
- Step 4: Burst all overflow requests exceeding local capacity to an OpenAI-compatible per-token API, scaling dynamically with zero idle waste.
Lyceum provides Serverless Inference to serve as the ideal elastic tier in this architecture. Running open-source models on European infrastructure, our platform delivers OpenAI-compatible endpoints with transparent per-token metering, with the current input and output rate for each model listed on its own catalogue record. By pairing reserved capacity for your baseline with per-token serverless inference for peak demand, you eliminate overprovisioning waste while maintaining high performance. Classify your traffic, then put the base load and the peak on different billing models.