Budget the task, not the request
A useful latency budget covers the complete task: model calls, tools, retrieval, queues and retries. Instrument the trace before choosing what to optimise. Model generation may dominate one workload while an external API dominates another.
A task-level budget is two ceilings, agreed in the product owner's terms before any architecture is fixed: how long a user will wait for the whole task, and what one completed task may cost. Anthropic's engineering guidance makes the underlying trade explicit: agentic systems often trade latency and cost for better task performance, so the trade is a choice you make, not a side effect you discover.
- Latency ceiling: the maximum wall-clock time a user waits for one completed task, set by the product, not the infrastructure.
- Cost ceiling: the maximum spend per completed task, covering every model call, tool call and retrieval hop in the run.
- Allocation: the split of both ceilings across step types, which is an engineering decision that comes after the ceilings exist.
Measure the distribution before setting limits
Trace a representative task sample. Record start and finish, each model or tool span, retries, queue time, token usage and the final outcome. Report median and 95th-percentile task latency alongside sample size. Also inspect the step-count distribution: a task with few slow calls can take longer than one with many fast calls. Do not multiply the 95th-percentile step count by the 95th-percentile step latency and call that a measured task percentile.
Allocate a worked task budget
The following is an illustrative planning budget, not a Lyceum benchmark. Suppose a task has a 20-second deadline and a $0.04 ceiling. Reserve $0.005 for retries, leaving $0.035 for the planned path. Model and tool costs below are assumed values; replace them with measured usage at the current rates.
| Stage | Latency allowance | Assumed cost |
|---|---|---|
| Plan | 3 seconds | $0.005 |
| Two independent lookups in parallel | 4 seconds | $0.004 total |
| Generate answer | 7 seconds | $0.018 |
| Validate result | 3 seconds | $0.006 |
| Queue and coordination allowance | 1 second | $0.002 |
| Reserve | 2 seconds | $0.005 |
The planned path totals 18 seconds and $0.035. Including the reserve reaches 20 seconds and $0.04. The lookup stage contributes the slower branch’s duration, plus coordination, while cost includes both calls. If the branches take 2 and 4 seconds, the parallel stage is about 4 seconds; sequential execution would take about 6 seconds. These are budget assumptions, not predictions of a latency percentile.
Parallelise independent work
Run calls together only when neither needs the other’s output. Bound concurrency so a fan-out cannot exhaust the request or spend budget. Cancel unnecessary work when supported, but account for calls that remain billable after cancellation. For dependent steps, improve the slow stage or reduce the number of round trips.
Enforce termination and preserve a useful result
Before starting another step, check remaining time, estimated maximum charge and in-flight reservations. For the illustrative task, stop starting new work after 18 seconds so 2 seconds remain to return the result. Cap retries and steps as additional limits. If the budget is exhausted, return completed findings, mark what remains unverified and offer a clear next action. Enforce these rules in code rather than asking the model to remember them.
Check quality and cost after the pilot
A cheap timeout is not a successful task. Report accepted tasks, failed tasks and capped tasks separately. Divide total spend across all attempts by the number accepted to obtain cost per success. Compare that with task latency percentiles and quality, then adjust limits using observed trade-offs. Do not claim an improvement until the revised workflow has been measured.