Serverless vs dedicated

Articles

30 September 2026

Self-Hosting vs a Managed API: The Break-Even Volume

The usual comparison divides a GPU hourly rate by theoretical throughput and calls self-hosting cheaper. Priced with utilisation and the full serving stack included, the crossover moves, and a managed dedicated endpoint sits between the two extremes.

11 September 2026

Self-Host vs API: The Token Break-Even for Open Models

If you are applying one self-hosting rule of thumb to every model you run, this guide shows how the crossover moves with model size and where it reverses entirely.

11 September 2026

Per-Token vs Per-GPU-Hour: Which Inference Pricing Fits

If you are choosing between paying per token and paying per GPU-hour, this guide reframes it as a question about who carries the utilisation risk and matches each option to a traffic shape.

1 June 2026

Open Source vs Closed API LLM Cost Comparison

API token prices have plummeted, but at scale, pay-as-you-go models still drain budgets. We work the arithmetic on where self-hosting open-source LLMs becomes cheaper than closed APIs, with every assumption shown.

20 May 2026

Inference Cost Per Token vs. Dedicated GPU: 2026 Economics

Token-based billing is a retail markup on compute. As your AI product scales, paying a US-based provider for every word generated becomes your largest line item. We break down the engineering math behind the switch to dedicated GPUs.

15 May 2026

LLM Inference Cost Per Token: Serverless vs. Dedicated Comparison

Inference cost per unit of model quality keeps falling, yet AI infrastructure bills continue to climb. We break down where dedicated GPUs become cheaper than serverless APIs, and how to work out your own threshold.

20 April 2026

Pay Per Token vs Dedicated GPU Inference: The Break-Even Guide

As hyperscaler credits expire, AI startups face a critical infrastructure fork: continue paying per token or move to dedicated GPUs. This guide breaks down the utilization math, latency trade-offs, and sovereignty requirements for European engineering teams.

Your next workload starts here