A pilot needs an endpoint, not a GPU

If evaluating an open coding model looks like an infrastructure project, this guide runs the pilot on a hosted endpoint and ends it with evidence instead of opinions. The usual blocker is the assumption that a pilot needs GPUs, a serving stack and someone to operate them. It does not. With a hosted OpenAI-compatible endpoint, setup is a configuration change in the editor, and the real work is deciding what to measure.

The question is not whether to evaluate, it is which stack. In the 2025 Stack Overflow Developer Survey, 84% of respondents reported using or planning to use AI tools in their development process, up from 76% the year before. For a team already paying per-seat licences for a closed coding assistant, that adoption curve is the pressure behind the pilot: the spend is visible, the renewal is coming, and the open-model alternative needs a fair test before procurement will sign anything.

Lyceum Serverless Inference hosts models behind a shared endpoint with per-token billing and no minimum commitment. You create an API key and configure the editor without provisioning a GPU. The steps below cover scope, criteria, endpoint, instrumentation, control and verdict. Collect enough observations for the decision you need, rather than assuming a fixed duration proves the result.

  1. Fix the scope: named participants, a fixed repository set, a two-week window, a defined task list.
  2. Agree success criteria and the exit decision before day one.
  3. Point the editor at a hosted endpoint: base URL, API key, model string.
  4. Instrument usage from day one: tokens, request volume, rejected completions.
  5. Randomly assign comparable tasks to the pilot and control, or counterbalance tool order.
  6. Read the result against the criteria, not against opinions.

Scoping the team, repos and tasks first

An unbounded pilot produces unbounded opinions. The moment participation is voluntary, the repository set is whatever people are working on, and the end date is whenever the budget line runs out, the result is a collection of anecdotes that procurement can neither accept nor reject. Fix four boundaries before installing anything:

  • Participants: start with a named pilot group, for example 5 to 8 people. This can reveal setup and usability problems; it does not guarantee a statistically reliable productivity estimate.
  • Repositories: a fixed set of codebases the team already knows well. Unfamiliar code confounds model quality with onboarding time.
  • Window: agree a calendar-dated period, for example 2 weeks, plus a rule for an inconclusive result. A short pilot may not capture learning effects or rare failures.
  • Tasks: a defined list of real work, written down before day one, so every participant attempts the same categories of task.

The METR randomized controlled trial is the scope pattern worth copying, and its design point is familiarity rather than sample size: 16 experienced developers completed 246 tasks in mature projects they had worked on for an average of five years, so the measurement isolated the tool rather than the learning curve. Pilot on repositories your team already knows for the same reason. The fuller reading of that study belongs in the control section below.

The security sign-off the scope needs

Before the first prompt leaves the building, the pilot scope needs a written answer to the data-protection question, because that answer determines whether the repositories in scope are eligible at all. On this hosted endpoint, prompts and outputs are processed, not stored, and never used for training; caching is in GPU memory only, per session, for minutes at most. A DPA carrying the named sub-processor list is available on request. That combination is what a works-council or data-protection review asks for first, and having it in the scope document before day one prevents the pilot stalling in week two while legal waits for a vendor reply.

Hosting is a per-model fact, not a platform-wide claim: check the individual model record for where the model you select runs, and state that region in the scope document exactly as the record states it. If your review requires processing inside EU jurisdiction, pick a model whose record says it is EU-hosted and write the region into the sign-off.

Agreeing success criteria before day one

Criteria agreed after the pilot ends are not criteria, they are arguments. Write them down before the first request, and mix two kinds of signal so neither can carry the verdict alone.

Measured criteria

  • Suggestion acceptance rate: record the share of suggestions kept, alongside correctness and review effort
  • Elapsed task time: compare randomly assigned matched tasks or a counterbalanced design, then check that quality meets the same standard

Reported criteria

A structured survey per participant, with fixed questions at fixed intervals, not an open feedback channel. The 2025 Stack Overflow survey found the biggest single frustration, cited by 66% of developers, is AI solutions that are almost right but not quite, followed by debugging AI-generated code being more time-consuming at 45%. Those are exactly the failure modes an open comment box buries under enthusiasm from the loudest participant. Ask instead: how often was a suggestion almost right, did fixing it cost more time than writing the code yourself, and would you keep the tool on this repository?

Write the exit decision now

Before day one, write down which result leads to adoption and which leads to stopping, and name the thresholds. A pilot that ends in adoption, a redesign, or a stop, decided in advance, produces a decision; a pilot that ends whenever the meeting happens produces a preference. For the structure of that decision document, including the four exit paths and the stop conditions, see AI Pilot Exit Criteria. One discipline matters above the others: claim no expected acceptance rate or productivity gain in advance. The pilot exists to measure them, and a forecast written into the criteria becomes the number the pilot is graded against regardless of what it measures.

Pointing the editor at a hosted endpoint

Which coding assistants can you point at your own OpenAI-compatible model endpoint? Most of the established ones, and the setup is the same three fields each time: a base URL, an API key, and a model string. This is the short part of the pilot. Every support statement below is dated to its documentation read, because tool support changes between releases.

  • Aider, checked 1 October 2026: set OPENAI_API_BASE and OPENAI_API_KEY, then use the openai/ prefix with the model ID
  • Cline, checked 1 October 2026: choose OpenAI Compatible and enter Base URL, API Key and Model ID in its settings
  • Roo, checked 1 October 2026: choose OpenAI Compatible, enter the connection fields and set model capabilities
  • Continue, checked 1 October 2026: use provider openai with apiBase, apiKey and model in YAML, and select chat completions for this endpoint
# Aider: configure the hosted chat-completions endpoint
export OPENAI_API_BASE="https://api.lyceum.technology/openai/v1"
export OPENAI_API_KEY="lk_your_api_key_here"
aider --model openai/deepseek/deepseek-v4-flash-0731

For the exact base URL and a worked walkthrough, the tool guides cover the three most common setups: How to Use Lyceum Models in Cline, the Cursor setup guide, and the Claude Code setup guide. Copy the base URL from the guide for your tool rather than retyping it; a trailing path fragment is the most common connection failure.

Features depend on the client and the model

Agent workflows need both a client that sends tool definitions and a model that returns usable tool calls. Roo documents native tool calling. Test it with your selected model and the actual file-editing workflow before the pilot. Cline’s browser or computer-use capabilities are separate from general function calling, so do not treat those labels as equivalent.

Claude Code uses the Anthropic protocol rather than OpenAI chat completions. Lyceum provides the separate route https://api.lyceum.technology/anthropic. Follow its current Claude Code setup guide for the required configuration instead of reusing the OpenAI endpoint.

Instrumenting usage from the start

Plan how to collect tokens, requests and errors before the pilot starts. API usage records and editor telemetry answer different questions, so confirm what each tool records and reconcile costs with billing. Instrumentation may need a script or exporter; it is not necessarily one config field.

  • Tokens per participant, input and output, per week.
  • Request volume per participant, so you can tell low usage from low value.
  • Failed or rejected completions: requests that errored, and suggestions the participant discarded. This is where the almost-right problem shows up in data.

Published per-model prices let you calculate pilot cost from measured input, cached-input and output usage. The invoice is the total to reconcile against, not automatically a per-participant log. Record usage by API key, project or editor session from day one. GLM-5.3-Flash lists $0.20 per million input tokens and $0.50 per million output tokens; check the current model record when you start.

Cline and Roo expose model-price fields for local cost estimates. Fill them from the provider’s price record, including cached rates where the client supports them. Collect each participant’s usage and errors separately, and reconcile the totals with the provider bill. A local estimate can differ from billing if prices, caching or token accounting differ.

A pilot without a control is a preference

Use a control that separates the tool’s effect from task familiarity. METR’s 2025 study randomly assigned 246 tasks among 16 experienced open-source developers to allow or disallow AI. In that setting, developers predicted faster work while measured completion time increased by 19%. That result is evidence from its particular tools, tasks and participants, not a prediction for your team or today’s models.

Avoid asking each participant to repeat the same task with both tools without accounting for the learning effect. Randomly assign comparable tasks or counterbalance tool order, record experience with each tool, and compare elapsed time with correctness and review effort. Treat participant impressions as a separate measure.

  • Use the same task categories, repositories and measurement window for both tools
  • Randomise comparable tasks or counterbalance tool order to reduce learning effects
  • Measure elapsed time, correctness, review effort and acceptance under both conditions

Reading the result against the criteria

Close the pilot by reading measurements against the criteria written before day one, not by collecting opinions afterward. The reading is four comparisons, each mapped to a criterion you already agreed:

  1. Suggestion acceptance rate: pilot tool versus control, per participant and pooled, over the full two weeks.
  2. Elapsed task time and quality: compare the pilot with the control using the agreed assignment or counterbalanced design.
  3. Tokens per participant against the published per-model prices: the cost line of the pilot, and the input to any seat-cost comparison procurement will ask for next.
  4. Structured survey scores: the reported signal, read alongside the measured ones, with the almost-right failure mode called out by name where it appears.

The decision that survives questioning names what was measured, over what period, against what control, and what result was agreed in advance to trigger adoption or stopping. It does not need to be favourable to the open model to be useful. A documented stop with a measured reason is a better procurement outcome than an undocumented adoption, because it preserves the option to re-run the pilot when the next model generation lands, and it took two weeks rather than a quarter.

The endpoint this pilot ran on is Lyceum Serverless Inference: pre-hosted open models behind one OpenAI-compatible endpoint, billed per token, nothing to provision. Agree the criteria, point the editor at a hosted endpoint, and measure from day one.