Budget per completed task, not per request
A coding task can involve many model and tool calls: planning, reading files, editing, testing and retries. Log the full run. Context handling depends on the harness, which may retain, trim, compact or restart the conversation.
Bai and colleagues studied eight frontier models on SWE-bench Verified in April 2026. They reported roughly 1,000 times the tokens of their code-chat and reasoning baselines, with input driving cost in that setup. Treat this as evidence of possible scale, not a production multiplier for every agent.
The fix is to restate the question. Not "what does a request cost" but "how many tokens does a completed task consume, of which kinds". That question has a defensible answer, and it converts directly into spend once you multiply by the per-token and cached-input rates of the model you run. The economics of that restatement, and why it usually favours open models billed per token over fixed seats, we cover separately in per-seat licences versus per-token inference.
- Measure every call in a task, including failed attempts and retries
- Split usage into fresh input, cached input and output at the applicable rates
- Report both cost per attempted task and total spend divided by accepted tasks
Where an agent's tokens actually go
Input volume can be large, but token count alone does not determine which component costs most. Cache discounts and output or reasoning prices can change the balance. Itemise these inputs before choosing what to reduce:
- Repository context: files, snippets and diffs included in each request
- Tool definitions: the schemas the client actually sends
- Tool results: their size and how long the harness retains them
- Conversation history: full, pruned or compacted according to the harness
- Retries: additional calls and any repeated context they carry
Context accumulates through the whole task
A harness that resends an expanding transcript can make later requests larger. Other harnesses trim tool results, compact history or start a new context. Inspect the requests your agent actually sends rather than assuming monotonic growth.
With equally sized additions and no pruning, total input across N steps can grow roughly with N squared. That is a simplified token-volume model, not a universal billing rule. Cache hits, variable output and compaction change the cost, and serving-cache memory is a separate measurement.
It also explains which tasks set your budget. The study found runs on the same task can differ by up to 30x in total tokens, and that accuracy often peaks at intermediate cost and saturates beyond it, meaning the most expensive runs are frequently unproductive exploration rather than deeper reasoning. The expensive tail of your task distribution, not the average, is what the budget has to survive. Two of the three levers below attack exactly that tail.
Caching the stable prefix
Put stable material before changing material when the provider supports prefix caching. A matching prefix creates an opportunity for a hit; routing, eviction and cache policy still determine whether it is reused. Use the returned cached-token usage to price the request.
Keep cache pricing and serving-engine implementation separate. A documented cache discount does not prove a particular engine or guarantee cross-request reuse. Check each provider’s minimum prefix size, retention options and usage fields; do not transfer one provider’s cache rules to another.
The Lyceum dashboard on 1 October 2026 listed Kimi-K2.7-Code at $1.25 fresh input and $0.31 cached input per million tokens. GLM-5.2 was $1.50 and $0.38 respectively. Both carried an EU label. Prices and advertised region are per-model observations.
| Model, dashboard checked 1 October 2026 | Fresh input per million | Cached input per million | Cached share of fresh price |
|---|---|---|---|
| moonshotai/kimi-k2.7-code | $1.25 | $0.31 | about 25% |
| z-ai/glm-5.2 | $1.50 | $0.38 | about 25% |
One data-residency note, because it matters to EU teams: prompts are cached in GPU memory only, per session, for a few minutes at most, and never written to a database. The discount does not require persisting your code.
Scoping context instead of handing over the repository
Caching reduces what the resent context costs. The second lever reduces what the context is. The worked example is aider's repository map. Instead of pasting whole files into the prompt, aider sends a concise map of the repository: the most important classes and functions with their types and call signatures, enough for the model to see how the code it is editing relates to the rest of the codebase.
The selection is not hand-tuned. Aider analyses a graph where each source file is a node and edges connect files with dependencies, applies a graph ranking algorithm, and selects the most relevant portions of the codebase to fit an active token budget set by the --map-tokens switch, which defaults to 1k tokens. The map also changes the interaction pattern: when the model needs more code, it uses the map to work out which files to look at and asks for those specific files, rather than the framework dumping the repository into context up front.
- Send a ranked summary: file list plus key symbols, fitted to a token budget.
- Let the model pull: it requests the specific files it needs, when it needs them.
- Effect per step: the stable prefix stays small, so both the cached and the uncached portion of every step shrink.
Capping iterations so failures terminate
Enforce a maximum step count, retry count and spend budget in the harness. A retry adds cost even when context has been compacted. Stop when the expected value of another attempt no longer justifies its cost, then return the partial result and failure reason.
- Maximum tool calls per task: a hard ceiling on loop length, enforced by the harness, not the model.
- Maximum retries per failing test: after N failed attempts on the same test, stop iterating on that approach.
- Token budget ceiling per task: terminate the run when total tokens cross the ceiling, whatever the model is mid-way through.
- Model routing per step: a fourth lever, sending the cheap steps to a cheaper model, which we treat separately in Cutting Inference Spend by Routing Requests by Difficulty.
Measuring a distribution of real tasks
Measure representative tasks from your own codebase and report the full distribution. Include failed and capped runs. Track the median and upper percentiles, plus the number of observations, so a few costly tasks remain visible.
Worked examples of that measurement exist in the literature. The SWE-agent paper reports the average cost per task for its own agent-computer interface setup on SWE-bench, and the Agentless paper does the same for its own pipeline. Treat both as examples of the method, never as transferable figures: each number is tied to one harness, one benchmark and one model, which is exactly why the only number that belongs in your budget is the one you measure on your own tasks.
- Input tokens and cached input tokens per task, from the API usage fields.
- Output tokens per task.
- Tool calls per task, and retries per failing test.
- End state: completed, capped, or escalated.
For each call, cost equals (fresh input × fresh rate + cached input × cached rate + output × output rate) / 1,000,000 when rates are per million tokens. Include billable reasoning in output without counting it twice. Sum all attempted runs and divide by accepted tasks for cost per success. If no task passes, that metric is undefined, not zero.