Public benchmarks do not cover your exact workload

SWE-bench contains 2,294 real repository issues from 12 Python projects, including changes across functions, classes and files. It is relevant evidence, but your languages, conventions and task mix may differ. A private evaluation tests that gap.

Historical changes provide a known starting commit and one reviewed solution. Build the instruction from the original issue or requirements, and keep the solution out of the model’s context. Public and private tasks both need controls against leakage and incomplete tests.

What actually fails

  • The agent loses track of an earlier edit and overwrites or contradicts a change it made two steps ago.
  • The agent emits an edit that does not apply: wrong context lines, a file path that does not exist, a hunk written against stale content.
  • The agent emits a malformed tool call, so the harness never receives a usable edit at all.

Patch application, build errors and malformed calls are mechanically checkable. They do not establish that the requested change is correct. Add task-specific acceptance checks and review any requirement that cannot be tested reliably.

One harness, one endpoint, many models

The evaluation only means something if the model is the only thing that changes between runs. Freeze everything else, and freeze it deliberately:

  • One agent harness, at one pinned version, for every model under test.
  • One set of tool definitions: the same file-read, file-edit and shell tools, described identically in every run.
  • One context budget: the same maximum context and the same repository snapshot logic, so no model sees more code than another.
  • One repository state: the same parent commit checked out for every task and every model.
  • One sampling configuration: temperature, top-p and token limits held fixed, even where a model card recommends different defaults.

Use the same harness where possible, but verify each model’s supported features and required settings. Record the endpoint, key scope, model identifier and any compatibility adjustments. One API schema does not guarantee identical tools or usage semantics.

Every major harness already supports this. Aider connects to any LLM behind an OpenAI-compatible endpoint by setting OPENAI_API_BASE and OPENAI_API_KEY and prefixing the model name with openai/. Roo Code and Cline take the same three settings in their provider panels: base URL, API key and model ID. Roo Code goes further and uses native tool calling exclusively, sending tool definitions to the model over the OpenAI tools schema. In your own harness, the OpenAI Python SDK accepts a custom base URL when the client is constructed, and the AI SDK exposes createOpenAICompatible for the same purpose.

One caution from the Roo Code documentation applies to every harness here: some OpenAI-compatible providers only partially implement the native tools API, so verify tool calling works before you trust a comparison run. A model that never receives well-formed tool events is being tested against a broken harness, not its own capability.

Reverting merged changes into test cases

The task set should come from code your team actually merged, not from a synthetic refactor. Start with the git log and look for merged changes that touched several files, ideally three or more, where the commit message states intent in a sentence or two. These exist in every repository with a few months of history, and they carry the property no synthetic task has: a human team once needed exactly this change.

  1. Choose a merged change with clear requirements and several affected files
  2. Check out its parent commit in a clean worktree
  3. Write an instruction from the original requirement, excluding the answer and implementation hints
  4. Prepare task-specific acceptance tests that fail on the parent and pass on the known good result
  5. Keep the reference diff outside the model context and record the commit, dependencies and all test commands

Provenance is what makes the set reproducible and auditable. Hugging Face makes the same argument for model cards: metadata such as license, base model and intended use is what makes a model discoverable and reproducible rather than a bare checkpoint. A task with a commit hash and a recorded repository state can be rerun a year later under the same starting conditions; a task described as the auth refactor cannot.

Treat the reference diff as review context, never as the pass criterion. It tells you and your client what a competent human did, which is useful when you read the runs afterwards. It must not decide pass or fail, for reasons the next section covers.

Making the build system the pass criterion

Require the following checks, and record each result separately. Existing tests passing on the untouched parent are a regression baseline, not evidence that the requested change happened:

  1. The patch applies to the recorded parent commit
  2. The project builds with the recorded dependencies and command
  3. Task-specific acceptance tests fail on the untouched parent and pass on the edited tree
  4. The existing regression suite still passes
  5. Any requirement outside automated coverage receives a recorded review

Similarity to the reference diff is a weak signal and is explicitly not the criterion. Many different edits can be correct: a rename in a different file, a helper extracted where the human inlined it, a guard clause instead of an early return. Scoring by diff similarity punishes a model for finding a second valid solution and rewards models that mimic surface shape. SWE-bench reaches the same conclusion from the other direction: it hands a model a codebase and an issue description and tasks it with editing the codebase to resolve the issue, evaluating the edited code rather than the shape of the patch.

Run the automated checks in continuous integration and emit a result per task and model. Include the acceptance-test precheck, so an untouched repository cannot receive a false pass. Keep review decisions separate from machine checks.

Recording four failure modes and the token bill

A single pass rate hides what to fix. Log four failure modes separately for every task and model:

  • Patch or tool failure: no usable edit was produced
  • Build failure: the edited project does not build
  • Acceptance or regression failure: the requested behaviour is missing or existing behaviour broke
  • Budget termination: the run exceeded its deadline, steps or spend limit

Failure categories guide investigation; they do not prove a cause. Inspect the trace before changing prompts or tool schemas. A build failure can have several causes, and another retry may not fix any of them.

Log fresh input, cached input and output, including billable reasoning and retries. Sum costs across every attempted run, then divide by accepted tasks. If none pass, cost per success is undefined. More steps do not imply a fixed cost multiplier because context, output and cache hits vary.

A full evaluation consists of independent task runs, but steps inside each agent run depend on earlier tool results. A static batch of prompts cannot reproduce that interaction. Batch only independent eligible requests; keep the interactive harness for dependent steps. No benchmark results are claimed here.

Too few tasks and drifting variables

State the sample size honestly. With ten tasks, a single task swings the pass rate by ten percentage points, so noise dominates any real difference between models. Aim for a set where one task moves the rate by a few points, which for most teams means several dozen tasks drawn from history. For scale, SWE-bench's full set runs to 2,294 problems; you do not need that to rank two models on your repository, but you do need enough tasks that a two-point difference is not one lucky patch.

The second threat is drift: variables that differ between runs and quietly invalidate the comparison. Pin each one in the run record:

VariablePin it as
Agent harnessName and exact version in the run record
Tool definitionsThe tool schema file, committed alongside the tasks
Context budgetMaximum context and repository snapshot rule, stated per run
Repository stateParent commit hash per task, checked out fresh
Sampling settingsTemperature, top-p and token limits, identical across models
Model identityAPI model string taken from the provider's model list on the day of the run

One drift lives on the provider side and is easy to overlook: the same open model can behave differently depending on how a provider serves it, so before you attribute a quality gap to the model, check whether your provider is quantizing the model. A comparison is only valid when every run's conditions are recoverable from the record.

Rerunning the set when a new model ships

When a new model becomes available, rerun the recorded tasks and compatibility checks. Keep the harness and acceptance criteria fixed where supported, and record necessary model-specific settings. A listed model identifier is discovery evidence; a successful test request confirms callability.

Record model metadata per run so results stay comparable over time: license, base model and parameter details, the fields Hugging Face documents as the metadata a model card should carry. Six months from now, a table without those columns cannot answer which version of which model produced which number.

This is where Serverless Inference fits the workflow: one OpenAI-compatible endpoint carrying the candidate models, billed per token with no GPU to provision, so the harness you built earlier is the only integration work. For the mechanics of pointing your existing SDK at an open model, the guide on switching the OpenAI SDK to an open-model endpoint covers the switch step by step. Build the task set from merged changes, then run every candidate model through the same endpoint and compare cost per completed task.