Benchmarks measure something narrower than your task
A benchmark can help build a shortlist, but inspect the task it actually measures. Some suites use short multiple-choice questions; others test multi-turn conversations, tools or repository changes. Match the benchmark’s task, scoring and limits to your product before treating its score as relevant.
- Dataset: identify the inputs, task types and difficulty distribution
- Execution: check turns, tools, context and time limits
- Scoring: distinguish exact-match accuracy, human preference and accepted task completion
- Operations: measure latency, failures and cost under your expected load
The original 2022 HELM study evaluated 30 models with multiple metrics across a common set of scenarios. Its lesson is methodological: measure more than accuracy and expose trade-offs. The original counts describe that study, not the current extent of HELM.
This is why a high score on a narrow suite transfers weakly to your failure modes. A model that aces multiple-choice reasoning benchmarks can still break your function-calling schema on the first production call, which is why open models with reliable function calling need to be tested on schema-bound outputs, not on abstract reasoning items. The benchmark measured a proxy. Your product is the task.
Who published it and who chose the comparison
The most common distortion is also the easiest to spot: the publisher chose the comparison set. A vendor announcing a new model picks its own baselines, its own prompt template, and its own scoring script, and the resulting table tells you what the publisher wanted to be compared against. A figure published by a model's own maker is vendor-reported, and any benchmark figure you cannot trace to an independent run should be attributed exactly that way, never restated as a measured fact.
It is worth knowing what a well-governed benchmark looks like, because the contrast is instructive. MLPerf Inference, run by MLCommons, publishes its rules openly, requires Closed division submissions to use the same model as the reference implementation so hardware comparisons stay apples-to-apples, identifies each submitter together with their software stack and code, and maintains a public change log for results that are later modified or invalidated. Every row is auditable. That is the standard a benchmark table should meet before it influences a procurement decision.
- Find who submitted the number: the model's own maker, or an independent party with nothing at stake in the ranking.
- Check whether the rules are public: dataset, prompt format, scoring script, and how ties or refusals were handled.
- Check whether the comparison set was fixed before results were known, or assembled after the winner was.
The practical check takes minutes. If you cannot find who ran the number and under what rules, treat the table as marketing collateral rather than evidence.
Configuration decides whether a number is comparable
A number without its configuration is not comparable to anything, including the same model scored somewhere else. The same checkpoint produces materially different scores depending on how it was served and prompted, and most benchmark tables you will meet in a launch post state none of it.
- Quantisation and checkpoint: record the exact artefact and precision
- Context: record input lengths, truncation and output budgets
- Sampling: disclose temperature, top_p, seeds and other supported controls
- Prompting: preserve system prompts, chat templates and examples
- Attempts: distinguish one-shot success from best-of-many results and include the cost of retries
The EleutherAI evaluation harness provides public task definitions and scoring implementations. Pin the harness version, task version and model settings before comparing results. A shared tool does not make runs comparable when those inputs differ.
Model cards are where this metadata belongs. The Hugging Face model-card documentation asks for training parameters, the datasets used, and evaluation results, so a reader can see what produced the number. If a benchmark claim comes with no configuration and no card entry backing it, you have an anecdote, not a measurement.
Contamination is hard to rule out from outside
Contamination can arise when evaluation material, or close variants, appears in training data. That can inflate scores without proving generalisation. Detection methods can identify some overlap or suspicious behaviour, but an external evaluator generally cannot conclusively rule out every exposure path.
The 2024 paper Leak, Cheat, Repeat reviewed 255 papers and reported about 4.7 million benchmark samples exposed to GPT-3.5 and GPT-4. The authors raised an indirect leakage risk under data-use policies. Exposure through use is not direct evidence that every sample entered a later training run or was memorised.
A new benchmark can still contain previously published material or leak after release. Treat provenance as part of evaluation quality. A published score is neither a guaranteed upper bound nor a prediction of production performance; contamination can bias it without establishing the size of that bias.
Where benchmarks are genuinely useful
Standardised runs within the same suite can support useful comparisons. In the original HELM study, a common evaluation setup increased coverage of the chosen scenarios. Read those historical results as evidence for a method, not as a current ranking or proof that every relevant failure mode is covered.
- Capability tiers: whether a small open model clears a rough quality bar at all, before you invest a day in a proper evaluation.
- Regression checks: whether a new checkpoint broke reasoning or instruction-following relative to its predecessor.
- Serving-system comparisons: MLPerf Inference, with its published rules, defined scenarios and metrics, and audited submissions, is the model of what a governed benchmark looks like for infrastructure claims.
Use benchmarks to narrow the candidates, then run a workload-specific pilot. Expand testing until the uncertainty is small enough for the decision, especially when failures are rare or costly.
Preference leaderboards and vendor figures have limits
The 2024 Chatbot Arena paper analysed more than 240,000 pairwise human votes. That is a historical study count, not a current platform total. Preference evaluation captures what sampled users preferred under those conditions, with uncertainty and sampling effects that a rank alone can hide.
But preference is not accuracy on your workload. Voters reward responses that read well: fluent, confident, well-formatted. Your product may need terse, schema-valid JSON on the first attempt, or a refusal to hallucinate a citation, and a model people enjoy reading can lose to a duller one on exactly those axes. The leaderboard answers what anonymous users preferred in a chat; your deployment answers a different question.
Vendor figures need the same separation. Before either kind of evidence reaches a deployment decision, sort your inputs: preference evidence tells you about reading experience, accuracy evidence tells you about task success, and the two should never be merged into one ranking column.
Start with a pilot, then size the evaluation to the decision
Start with 30 examples if you need a quick failure-finding pilot. That sample is not enough for every release decision. Use a representative held-out sample for aggregate quality, and a separate targeted set for known edge cases. If you tune prompts on the pilot, evaluate the final choice on fresh held-out examples.
- Sample permitted production cases across the traffic types you need to support.
- Keep representative cases separate from targeted failure tests.
- Define the scoring rubric and acceptance criteria before comparing candidates.
- Record each model’s supported settings, output budget and repeated-run policy.
- Review failures, report uncertainty and expand the sample where the decision remains unclear.
As an illustrative token calculation, 30 calls with 1,000 input and 500 output tokens each use 30,000 input and 15,000 output tokens. At Qwen3.5-9B’s listed $0.15 input and $0.20 output rates, that is $0.0075 per run. Longer prompts, reasoning, retries and repeated trials increase the bill. The rates were checked on 1 October 2026; count all billable tokens.
Use published results to choose candidates, then test the actual endpoints on your workload. Lyceum’s serverless API lets you evaluate pre-hosted models with token billing. Configure the base URL, key and model, and include evaluation charges in your budget.