Find the requests that a smaller model can handle
If every request in your product goes to the strongest model, this guide segments your traffic by difficulty and builds the quality gate that makes routing safe. The pattern is everywhere: a prototype was built on the strongest model available, the prototype shipped, and a year later classification calls, JSON extraction and short factual lookups are all billed at frontier prices. Nobody chose that spend. It is the residue of a model string that was never revisited.
Classification, extraction and formatting are candidates for a smaller model, not guarantees of equal quality. Complex schemas, unusual inputs and domain knowledge can make a short request difficult. Compare current input, output and cache prices for the exact models.
| Request shape | Quality checks | Possible routing signal |
|---|---|---|
| Classification | Accuracy by class and rare cases | Known task type plus calibrated quality evidence |
| Extraction | Schema and field correctness | Schema checks and task-specific validation |
| Rewrite or translation | Meaning, style and omissions | Length, language and evaluated task segment |
| Factual answer | Evidence and citation accuracy | Retrieval quality and verification |
| Code or reasoning | Acceptance tests and correctness | Task segment, tests and measured failure rate |
Sampling and segmenting real traffic first
You cannot route what you have not measured. The failure mode we see most often is a team picking a router before looking at a single logged request, then discovering the router optimises a mix that does not exist. Start with a sample of real production traffic, large enough to cover weekly seasonality, and classify each request by what it requires: the task type, the output format it must honour, and how much latency it tolerates. Not by the endpoint it arrived at. An internal endpoint name tells you which team shipped the feature, nothing about difficulty.
- Task type: classification, extraction, formatting, factual lookup, or open-ended generation.
- Output contract: free text, strict JSON, or a schema with required fields.
- Latency tolerance: interactive, or can it wait minutes or hours.
- Volume: requests per day, so the saving math has a denominator.
Independent offline requests can be candidates for discounted batch processing when the selected service supports it. Confirm the model, rate, completion window and quality settings. Dependent agent steps cannot simply be sent as a static batch.
Build a segment table with volume, task, acceptable latency and current model. The following allocation is illustrative, using dashboard prices checked on 1 October 2026.
| Segment | Latency | Currently served by | Input / output per 1M tokens |
|---|---|---|---|
| Ticket triage (label from 12 options) | Minutes matter little | GLM-5.3 | $1.40 / $4.40 |
| Field extraction to strict JSON | Interactive | GLM-5.3 | $1.40 / $4.40 |
| Contract clause summarisation | Hours acceptable | Kimi-K3 | $3.00 / $15.00 |
| Multi-step code reasoning | Interactive | Kimi-K3 | $3.00 / $15.00 |
The table is an illustrative workload allocation using model prices, not a measured customer result. A cheaper candidate is useful only if it meets the segment’s quality and latency requirements.
Three routing signals and what each costs
A routing signal is anything that decides, before or instead of the expensive call, which model serves a request. There are three useful ones, ordered by how much they cost to build and run. Pick the cheapest one that separates your traffic, and accept that published benchmark numbers belong to their benchmarks, not to your task mix.
| Signal | How it decides | Cost to run | Catches |
|---|---|---|---|
| Static rules | Request type, endpoint, prompt length, a flag in the payload | Effectively zero, a dict lookup | Only difficulty visible in the request itself |
| Learned router | A small classifier scores each prompt and routes by that score; RouteLLM's routers cut GPT-4 calls to 14% of traffic while keeping 95% of GPT-4 performance on MT Bench | One small forward pass per request | Difficulty that correlates with prompt features |
| Cascade with escalation | Generate, verify, then escalate if needed | Cheap generation + verification + any escalation | Failures detected by the chosen checks |
Trying the cheap model first and escalating
The cascade's mechanics are unglamorous, which is why it works. On Lyceum Serverless Inference the small and the large model sit behind the same OpenAI-compatible endpoint, so a cascade changes only the model string between the cheap call and the escalated call, and each call is billed at its own published per-token price. No second client, no second auth path, no infrastructure to run.
This minimal Python sketch validates a verifier’s numeric output and escalates on an invalid score. Supply a threshold chosen on calibration data. It is not a calibrated production router; retain task-specific checks and a separate held-out quality test.
import math
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.lyceum.technology/openai/v1",
api_key=os.environ["LYCEUM_API_KEY"],
timeout=30.0,
max_retries=0,
)
CHEAP = "z-ai/glm-5.3-flash"
LARGE = "z-ai/glm-5.3"
def parse_score(value):
try:
score = float(value)
except (TypeError, ValueError):
return None
return score if math.isfinite(score) and 0 <= score <= 1 else None
def cascade(messages, *, threshold):
if parse_score(threshold) is None:
raise ValueError("threshold must be finite and between 0 and 1")
threshold = float(threshold)
calls = []
def ask(model, prompt):
response = client.chat.completions.create(
model=model, messages=prompt, max_tokens=4096
)
calls.append({"model": model, "usage": response.usage})
return response.choices[0].message.content or ""
draft = ask(CHEAP, messages)
score = None
if draft.strip():
check = ask(CHEAP, messages + [
{"role": "assistant", "content": draft},
{"role": "user", "content":
"Assess the preceding answer. Return only a score from 0 to 1."},
])
score = parse_score(check)
# A score is a routing signal, not a probability of correctness.
if score is not None and score >= threshold:
return {"answer": draft, "path": "cheap", "calls": calls}
answer = ask(LARGE, messages)
return {"answer": answer, "path": "escalated", "calls": calls}
Let C_g be cheap generation cost, C_v verification cost and C_r any routing overhead. If p is the share accepted without escalation and C_l is the large-call cost, expected cost is C_g + C_v + C_r + (1 − p) × C_l under constant per-path averages. It beats a large-only cost C_l when p > (C_g + C_v + C_r) / C_l. Measure those costs from actual token usage; a ratio of list prices alone is insufficient. Check region and data handling for every model in both paths.
For illustration, if generation costs $0.002, verification $0.001 and escalation $0.02, with no other routing charge, the break-even acceptance share is 15%. At 60% acceptance, expected cost is $0.003 + 0.40 × $0.02 = $0.011, versus $0.02 large-only. These assumed figures do not establish quality or measured savings.
Building the quality gate before the switch
A self-reported score is not a calibrated probability. Label representative requests, choose the threshold on calibration data, then evaluate the frozen router on a separate held-out set. Measure wrong answers accepted by the cheap path as well as unnecessary escalations.
- Build a held-out set per segment, drawn from real production requests, sized so the metric you care about is stable.
- Score every item on both paths: the current model and the proposed cheap model.
- Agree the accepted difference per segment in writing before the first score is computed.
- Route only the segments that pass. A segment that fails stays on the large model, full stop.
- Record the passing scores as the baseline the monitoring section tracks against.
When escalation costs more than not routing
A strict threshold can raise escalation cost; a loose one can pass bad answers. Compare actual total generation, verification, routing and escalation charges with the large-only baseline. If call lengths differ by route, use measured conditional costs rather than the simplified constant-cost formula.
- Threshold too tight: escalation rate too high, expected cost crosses above C_large, and routing loses money. Watch the escalation rate against the break-even p.
- Gate too loose: the cascade never escalates, the dashboard shows a beautiful saving, and silent quality loss ships to users. The held-out scoring from the quality gate is what catches this.
- Latency compounding: every escalation adds the cheap call's time, plus the verifier's, to the request. A latency-critical interactive segment may belong on static rules or on the large model outright, even if it would pass the quality gate.
Monitoring as the traffic mix shifts
A routing rule is fitted to the traffic you sampled, and traffic mixes shift. A feature launch moves volume between segments, a prompt template change alters the difficulty distribution, and a rule fitted to last quarter degrades quietly, because nothing crashes when a router starts making worse decisions. Monitoring is the verify step, and it has three parts: re-score the held-out set per segment on a schedule, track the escalation rate against the break-even p, and track cost per segment against the baseline recorded at the quality gate.
- Re-score the held-out set per segment on a fixed schedule, comparing against the baseline scores from the gate.
- Track escalation rate and cost per segment; a drifting escalation rate is the earliest sign the mix has moved.
- Re-read per-token prices from each model's own record when the mix shifts; the model roster changes, and Lyceum Model Roster, Prices and Changes tracks those updates.
- Treat published router results as priors, not guarantees. RouteLLM reports its routers maintain performance when the strong and weak models are changed at test time, but your own mix is the only benchmark that counts.
Score both paths on a held-out set per segment, then route only the segments that pass. Fill the break-even condition with your own prices from the live per-token pricing page, and the segments that pass are the ones that stop paying frontier prices.