Provider budget alerts arrive too late

The first signal that LLM spend is out of control is usually the invoice, because the controls most teams reach for first are the provider's own. OpenAI lets you configure a spend alert, which sends a notification while API traffic continues, and a hard spend limit, at which point affected API requests return a 429 error. Both are enforced at the organization or project level. If LLM spend across your team is only visible on the invoice, this guide puts attribution and an enforceable limit at the layer where they actually work. That matters more as developer AI spend becomes a visible line item and teams replace per-seat licences with open models billed per token, a shift we analysed in per-seat licences versus per-token inference.

Anthropic draws the same distinction in its own documentation: spend limits set a maximum monthly cost an organization can incur, while rate limits set the maximum number of API requests over a defined period. Both are enforced at the organization level, though you can set lower user-configurable limits per workspace.

A shared credential makes attribution harder, but providers may expose projects, service accounts and separate keys. Use those controls where they meet your requirements. Add a gateway when you need shared policy across providers or finer account mapping.

Provider controlWhat happens at the limitWhere it applies
OpenAI spend alertSends a notification; traffic continuesOrganization or project
OpenAI hard spend limitAffected requests return a 429Organization or project
Anthropic spend capUsage pauses until the first day of the next month; requests return HTTP 429Organization, with optional lower workspace limits
Anthropic rate limitsRequests per minute and tokens per minute, per model classOrganization

Issuing a key per team and per service

Before any limit can be enforced, every request has to carry an identity the control layer can read. This is the attribution prerequisite: without a key or header per team and per service, no limit can be enforced against the right caller and no report can be produced after the fact. Get this wrong and every later step is guesswork.

The mechanics are well established. LiteLLM's proxy issues virtual keys through a /key/generate endpoint, and spend is tracked automatically per key, per user, and per team in its database tables. A key is not just a credential; it carries policy:

  • Assign a clear owner and environment to each virtual key
  • Configure and test explicit key and team budgets rather than assuming inheritance
  • Check the deployed release’s reporting endpoints and reconcile their totals with provider usage

Budget enforcement depends on the gateway’s stored counters, request accounting and behaviour during faults. For your deployed version, test concurrent requests, restarts and unavailable storage. Confirm whether requests fail closed or continue, and document any permitted overshoot. Do not treat a configured number as proof of enforcement.

Putting a gateway in the request path

Controls can sit at the provider, gateway and application. Provider projects can isolate spend and attribute usage; a gateway can combine multiple providers; the application can enforce a task-specific budget. Choose complementary controls and test their shared limits.

LayerAttributionBlocking scope
ProviderProjects, service accounts or keys, where supportedEnforced account or project limits, depending on provider
GatewayMapped caller, team, service or environmentConfigured key or shared budgets before forwarding, subject to accounting delay
ApplicationUser, feature or taskTask quotas, iteration limits and deadlines

Lyceum’s chat endpoint is https://api.lyceum.technology/openai/v1. A gateway can route to an OpenAI-compatible endpoint, but test streaming, tool calls, usage accounting and errors before using its budgets. Configure the provider key and model identifier as well as the base URL.

In practice this means one control plane for the whole team: the same virtual keys, the same budgets and the same reports cover every provider the gateway fronts, instead of a separate alert configuration per vendor dashboard.

Setting a soft alert and a hard limit

A spend control needs two settings, not one. A soft alert informs: it fires when a team approaches its ceiling, while the service keeps running and someone still has time to react. A hard limit refuses: at the ceiling, the gateway rejects the request. Teams need both, because an alert alone cannot stop spend and a hard limit alone gives no warning before it takes a service down.

For a gateway budget, record its scope, amount, reset rule and action at the limit. Confirm which key and team policies apply in your deployed release. Test those controls with a small quota and simultaneous calls. Task iteration limits belong in the agent harness unless your chosen gateway explicitly supports them.

OpenAI documents both alerts and enforced organisation or project hard spend limits. A hard limit can return HTTP 429, although accounting delays can allow overshoot. Map each team or service to the available provider scopes, or enforce additional scopes through your application or gateway.

Rate limits catch what spend caps cannot

A hard spend cap can stop aggregate loop spend when it takes effect. It cannot stop a loop from rapidly using the allowed budget first. Rate limits slow request or token volume, while a task deadline and iteration cap terminate the loop itself.

Providers meter both dimensions. OpenAI's rate limits use RPM (requests per minute), RPD (requests per day), TPM (tokens per minute) and TPD (tokens per day), and a request counts against whichever limit it hits first. Anthropic measures requests per minute (RPM), input tokens per minute (ITPM) and output tokens per minute (OTPM) per model class, enforced with a token bucket that replenishes continuously rather than resetting at fixed intervals.

At the gateway, Kong's AI Rate Limiting Advanced plugin extends this to token-aware limiting: it can rate limit on total, prompt or completion tokens, or on computed cost, calculated as (prompt_tokens × input_cost + completion_tokens × output_cost) / 1,000,000. Policies can match on consumer, consumer group, header, path, model or provider, and when a limit is reached the plugin returns an HTTP 429 with headers showing the limit, the remaining quota and the reset time.

One note on where the provider's own limit sits. Some providers size rate limits to the customer's traffic rather than to fixed tiers, so the gateway's rate limit is your own control to set against your real traffic shape, not a provider tier you have to engineer around. Set it per key: a production service gets headroom, an experiment gets a fraction of it.

Separating experimentation from production budgets

The most common failure after attribution is in place is a shared budget pool. When experiments and production draw on the same ceiling, a single evaluation run can exhaust the pool and the hard limit takes a customer-facing service down with it. The fix is to split budgets by environment: a key for development, a key for staging, a key for production, each with its own soft alert and hard limit, so an experiment can only degrade its own budget.

  • Use separate budgets for development, staging and production, and check shared organisation limits
  • Send only eligible independent offline work to batch processing
  • Meter queued batch work separately and verify cancellation, completion and billing behaviour

Batch is also the cheaper budget to run experiments in: async batch inference is billed at a discount to the list price, so offline evaluations sit in a separate and materially smaller budget than the production endpoint. We cover the trade-offs in batch versus real-time inference pricing.

Assign an owner to each budget and review it as the workload matures. Include the time and operational cost of maintaining a gateway when deciding whether provider controls are enough.

Checking the report attributes every request

Once keys, gateway and limits are in place, the verification step is to read the report and confirm it answers the questions a spend cap cannot: not only what the month cost, but which service spent it, on what, and whether that spend was necessary. Three metrics do most of the work:

  • Tokens per request per service: catches prompt bloat and a service that quietly sends entire documents where a summary would do
  • Requests per service per day: catches loops and retry storms before they reach a monthly ceiling
  • Model per service: catches an expensive default, so a service calling a premium model for a summarisation task can be switched to a cheaper one

Report spend by service, model and environment. Reconcile fresh, cached and output token charges with provider invoices. A blended cost per million tokens changes with workload mix, so it cannot by itself establish which provider is cheaper.

This is where per-token billing makes the whole exercise meaningful. A limit expressed in dollars only binds if you can read the price of a request directly off the rate card. Lyceum's Serverless Inference meters pre-hosted open models per token with no minimum commitment, and publishes exact per-token prices per model on its live per-token pricing page. Issue a key per service, put a gateway in the path, then set the soft and hard limits against per-token prices you can read.