Find the requests that a smaller model can handle

If every request in your product goes to the strongest model, this guide segments your traffic by difficulty and builds the quality gate that makes routing safe. The pattern is everywhere: a prototype was built on the strongest model available, the prototype shipped, and a year later classification calls, JSON extraction and short factual lookups are all billed at frontier prices. Nobody chose that spend. It is the residue of a model string that was never revisited.

Classification, extraction and formatting are candidates for a smaller model, not guarantees of equal quality. Complex schemas, unusual inputs and domain knowledge can make a short request difficult. Compare current input, output and cache prices for the exact models.

Request shapeQuality checksPossible routing signal
ClassificationAccuracy by class and rare casesKnown task type plus calibrated quality evidence
ExtractionSchema and field correctnessSchema checks and task-specific validation
Rewrite or translationMeaning, style and omissionsLength, language and evaluated task segment
Factual answerEvidence and citation accuracyRetrieval quality and verification
Code or reasoningAcceptance tests and correctnessTask segment, tests and measured failure rate

Sampling and segmenting real traffic first

You cannot route what you have not measured. The failure mode we see most often is a team picking a router before looking at a single logged request, then discovering the router optimises a mix that does not exist. Start with a sample of real production traffic, large enough to cover weekly seasonality, and classify each request by what it requires: the task type, the output format it must honour, and how much latency it tolerates. Not by the endpoint it arrived at. An internal endpoint name tells you which team shipped the feature, nothing about difficulty.

  • Task type: classification, extraction, formatting, factual lookup, or open-ended generation.
  • Output contract: free text, strict JSON, or a schema with required fields.
  • Latency tolerance: interactive, or can it wait minutes or hours.
  • Volume: requests per day, so the saving math has a denominator.

Independent offline requests can be candidates for discounted batch processing when the selected service supports it. Confirm the model, rate, completion window and quality settings. Dependent agent steps cannot simply be sent as a static batch.

Build a segment table with volume, task, acceptable latency and current model. The following allocation is illustrative, using dashboard prices checked on 1 October 2026.

SegmentLatencyCurrently served byInput / output per 1M tokens
Ticket triage (label from 12 options)Minutes matter littleGLM-5.3$1.40 / $4.40
Field extraction to strict JSONInteractiveGLM-5.3$1.40 / $4.40
Contract clause summarisationHours acceptableKimi-K3$3.00 / $15.00
Multi-step code reasoningInteractiveKimi-K3$3.00 / $15.00

The table is an illustrative workload allocation using model prices, not a measured customer result. A cheaper candidate is useful only if it meets the segment’s quality and latency requirements.

Three routing signals and what each costs

A routing signal is anything that decides, before or instead of the expensive call, which model serves a request. There are three useful ones, ordered by how much they cost to build and run. Pick the cheapest one that separates your traffic, and accept that published benchmark numbers belong to their benchmarks, not to your task mix.

SignalHow it decidesCost to runCatches
Static rulesRequest type, endpoint, prompt length, a flag in the payloadEffectively zero, a dict lookupOnly difficulty visible in the request itself
Learned routerA small classifier scores each prompt and routes by that score; RouteLLM's routers cut GPT-4 calls to 14% of traffic while keeping 95% of GPT-4 performance on MT BenchOne small forward pass per requestDifficulty that correlates with prompt features
Cascade with escalationGenerate, verify, then escalate if neededCheap generation + verification + any escalationFailures detected by the chosen checks

Trying the cheap model first and escalating

The cascade's mechanics are unglamorous, which is why it works. On Lyceum Serverless Inference the small and the large model sit behind the same OpenAI-compatible endpoint, so a cascade changes only the model string between the cheap call and the escalated call, and each call is billed at its own published per-token price. No second client, no second auth path, no infrastructure to run.

This minimal Python sketch validates a verifier’s numeric output and escalates on an invalid score. Supply a threshold chosen on calibration data. It is not a calibrated production router; retain task-specific checks and a separate held-out quality test.

import math
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.lyceum.technology/openai/v1",
    api_key=os.environ["LYCEUM_API_KEY"],
    timeout=30.0,
    max_retries=0,
)
CHEAP = "z-ai/glm-5.3-flash"
LARGE = "z-ai/glm-5.3"


def parse_score(value):
    try:
        score = float(value)
    except (TypeError, ValueError):
        return None
    return score if math.isfinite(score) and 0 <= score <= 1 else None


def cascade(messages, *, threshold):
    if parse_score(threshold) is None:
        raise ValueError("threshold must be finite and between 0 and 1")
    threshold = float(threshold)
    calls = []

    def ask(model, prompt):
        response = client.chat.completions.create(
            model=model, messages=prompt, max_tokens=4096
        )
        calls.append({"model": model, "usage": response.usage})
        return response.choices[0].message.content or ""

    draft = ask(CHEAP, messages)
    score = None
    if draft.strip():
        check = ask(CHEAP, messages + [
            {"role": "assistant", "content": draft},
            {"role": "user", "content":
             "Assess the preceding answer. Return only a score from 0 to 1."},
        ])
        score = parse_score(check)
    # A score is a routing signal, not a probability of correctness.
    if score is not None and score >= threshold:
        return {"answer": draft, "path": "cheap", "calls": calls}
    answer = ask(LARGE, messages)
    return {"answer": answer, "path": "escalated", "calls": calls}

Let C_g be cheap generation cost, C_v verification cost and C_r any routing overhead. If p is the share accepted without escalation and C_l is the large-call cost, expected cost is C_g + C_v + C_r + (1 − p) × C_l under constant per-path averages. It beats a large-only cost C_l when p > (C_g + C_v + C_r) / C_l. Measure those costs from actual token usage; a ratio of list prices alone is insufficient. Check region and data handling for every model in both paths.

For illustration, if generation costs $0.002, verification $0.001 and escalation $0.02, with no other routing charge, the break-even acceptance share is 15%. At 60% acceptance, expected cost is $0.003 + 0.40 × $0.02 = $0.011, versus $0.02 large-only. These assumed figures do not establish quality or measured savings.

Building the quality gate before the switch

A self-reported score is not a calibrated probability. Label representative requests, choose the threshold on calibration data, then evaluate the frozen router on a separate held-out set. Measure wrong answers accepted by the cheap path as well as unnecessary escalations.

  • Build a held-out set per segment, drawn from real production requests, sized so the metric you care about is stable.
  • Score every item on both paths: the current model and the proposed cheap model.
  • Agree the accepted difference per segment in writing before the first score is computed.
  • Route only the segments that pass. A segment that fails stays on the large model, full stop.
  • Record the passing scores as the baseline the monitoring section tracks against.

When escalation costs more than not routing

A strict threshold can raise escalation cost; a loose one can pass bad answers. Compare actual total generation, verification, routing and escalation charges with the large-only baseline. If call lengths differ by route, use measured conditional costs rather than the simplified constant-cost formula.

  • Threshold too tight: escalation rate too high, expected cost crosses above C_large, and routing loses money. Watch the escalation rate against the break-even p.
  • Gate too loose: the cascade never escalates, the dashboard shows a beautiful saving, and silent quality loss ships to users. The held-out scoring from the quality gate is what catches this.
  • Latency compounding: every escalation adds the cheap call's time, plus the verifier's, to the request. A latency-critical interactive segment may belong on static rules or on the large model outright, even if it would pass the quality gate.

Monitoring as the traffic mix shifts

A routing rule is fitted to the traffic you sampled, and traffic mixes shift. A feature launch moves volume between segments, a prompt template change alters the difficulty distribution, and a rule fitted to last quarter degrades quietly, because nothing crashes when a router starts making worse decisions. Monitoring is the verify step, and it has three parts: re-score the held-out set per segment on a schedule, track the escalation rate against the break-even p, and track cost per segment against the baseline recorded at the quality gate.

  • Re-score the held-out set per segment on a fixed schedule, comparing against the baseline scores from the gate.
  • Track escalation rate and cost per segment; a drifting escalation rate is the earliest sign the mix has moved.
  • Re-read per-token prices from each model's own record when the mix shifts; the model roster changes, and Lyceum Model Roster, Prices and Changes tracks those updates.
  • Treat published router results as priors, not guarantees. RouteLLM reports its routers maintain performance when the strong and weak models are changed at test time, but your own mix is the only benchmark that counts.

Score both paths on a held-out set per segment, then route only the segments that pass. Fill the break-even condition with your own prices from the live per-token pricing page, and the segments that pass are the ones that stop paying frontier prices.