The API is the cheapest part of a switch

Your product runs on inference, your current provider has given you one more reason to leave, and the migration estimate circulating internally says a few days of work because the API is OpenAI-compatible. That estimate prices the cheapest of five cost components and silently drops the other four. If you are weighing a move to another inference provider, this guide costs the five components properly and shows which one usually dominates.

For a basic supported OpenAI-style chat call, the initial configuration change is usually the base URL, API key and model identifier. Your streaming parser, tool handling, retry policy and usage accounting may also need changes. Test them against the new endpoint.

  • Base URL: point the SDK client at the new provider's OpenAI-compatible base URL
  • API key: replace the key, one placeholder
  • Model string: swap in the target model's public identifier, for example deepseek/deepseek-v4-flash-0731
  • Still different: tool-calling schemas, JSON mode behaviour, max token handling, stop sequence behaviour
  • Still different: error shapes, rate-limit headers, and how each provider documents the above

A shared request format does not guarantee shared behaviour. Different checkpoints, quantisation, chat templates, defaults and safety handling can change outputs. Compare the actual endpoints on your held-out tasks, even when the public model names match.

Start with the configuration change, then test every feature used by your application. A small code diff can still require substantial validation. The migration estimate should include that work alongside any parser or parameter changes.

What a switch is actually worth to you

Estimate 5 components separately. Their relative cost depends on your application and existing tests:

  1. Output-quality re-validation: building and running an evaluation set from your own traffic to prove the new provider is not worse
  2. Prompt re-tuning: re-fitting prompts that were tuned against the old model's formatting behaviour
  3. Integration and parameter differences: tool calling, JSON mode, max tokens, stop behaviour, error handling
  4. Parallel running: paying two providers for the same traffic while you build up trust in the new one
  5. The organisational cost of the decision: the review, the sign-off and the risk discussion that precedes any of the engineering

A team with a mature evaluation suite may spend more on parallel traffic, while a team without one may spend more building tests. There is no universal ranking or duration. Estimate engineering days and API charges for your own stack, and record the assumptions.

Treat evaluation as part of the migration itself. Keeping the same model may narrow the change; choosing different weights can widen it. In either case, test the behaviour users depend on.

Costing output re-validation first

A held-out evaluation set gives the team a defined basis for accepting the new endpoint. Include its creation and review in the estimate, and keep the final acceptance cases separate from prompt-tuning examples.

  • Sample real production prompts, not benchmark data, so the distribution matches what your users send
  • Hold the set out: it must never have been used to fit the current prompts, or it will flatter them
  • Label expected behaviour for each case, including tool calls, structured output and refusal handling
  • Cover the edges your product actually hits: long context, multilingual input, adversarial formatting, streaming interruptions
  • Decide the pass threshold before the first run, so the result cannot be argued after the fact

An existing evaluation framework can help run and score cases, but it does not make production prompts private by itself. Review where the harness sends inputs, stores logs and uploads results. Use permitted data and the same processing requirements as the production system.

Price the runs themselves. Every evaluation pass consumes API tokens on both providers, and a serious suite gets run more than once: once to baseline the incumbent, once per candidate, again after prompt re-tuning. Those token costs belong in the switch estimate as a line item, not in the noise.

Prompt re-tuning and parameter differences

Prompt formatting can affect model output. Keep a baseline with the existing prompt, then test changes on a tuning set if needed. Do not assume every switch requires a rewrite, or that a prompt tuned for one endpoint will transfer without checking.

Budget prompt work separately from API integration. Compare candidate prompts on a tuning set, then use held-out cases for the acceptance decision. Repeatedly tuning against the final test set makes its reported quality less informative.

Parameter surfaceWhat differs between providersWhat to check before cutover
Tool callingSchema binding, whether the model emits tool calls natively or in text, parallel call handlingRun your full tool registry against the evaluation set; verify constrained decoding behaves
JSON modeWhether JSON mode is enforced at the engine level or requested in the prompt, malformed-output rateCount unparseable responses on the same prompt set on both providers
Max tokensWhether limits count prompt plus completion, truncation behaviour at the ceilingConfirm the limit semantics match your accounting before load testing
Stop sequencesWhether stop strings are stripped from output, how multiple sequences interactDiff raw completions against your current provider on stop-heavy prompts
StreamingChunk granularity, time-to-first-token, keep-alive behaviour under loadMeasure p95 time-to-first-token under production-shaped concurrency

Tool calling and JSON mode deserve the most attention, because they are where a compatible API is least compatible in practice. If your product depends on structured output, pick from open models with reliable function calling and JSON output and verify the behaviour empirically before the shadow phase, not after.

Price parallel traffic at the new provider’s rates

Mirroring requests adds the new provider’s cost while the original production traffic continues. The total does not necessarily double: prices, tokenisation, output lengths and cache hits can differ. Estimate the incremental bill from the mirrored workload.

  1. Estimate the mirrored input, cached-input and output tokens for the actual shadow period.
  2. Multiply each token category by the new provider’s applicable rate per million.
  3. Add retries, reasoning output and other billed request types.
  4. Add separate evaluation runs and storage or gateway charges.
  5. Add that incremental amount to the incumbent bill for the same period.

For example, 10 million uncached input tokens and 2 million output tokens at DeepSeek-V4-Flash-0731’s listed Lyceum rates cost $2.50 + $0.60 = $3.10. Those rates were checked on 1 October 2026. If an eligible asynchronous batch job is billed at half those rates, the same mix is $1.55. Offline batch replay does not measure live latency or reproduce an interactive tool loop.

Choose the shadow duration and sample from the risks you need to test. Some systems can use offline replay followed by a small canary. A full duplicate bill is not an unavoidable migration cost, and a fixed number of weeks does not prove sufficient coverage.

What keeping the same model can reduce

Keeping the same checkpoint can reduce the number of changes under test. It does not make quality identical. Providers can use different quantisation, templates, sampling defaults, context handling or serving settings. Retain task-level acceptance tests alongside infrastructure checks.

  • Latency: p95 time-to-first-token and end-to-end latency under production-shaped concurrency
  • Throughput: sustained tokens per second at your real concurrency, not a benchmark figure
  • Rate limits: the ceiling your traffic actually hits at peak, and what happens when you reach it
  • Quantisation: whether the served weights are full precision or quantised, because quantisation changes output; how to tell whether a provider quantizes the model covers the checks
  • Context limits: the effective context window against your longest production prompts

Confirm the exact checkpoint and serving configuration where the provider exposes them. A public model identifier alone does not prove that both endpoints serve an identical artefact. Run quality, tool-use and output-format tests as well as load tests.

If no provider serves the weights you want, the alternative is hosting them yourself, and that trade has its own arithmetic: the token break-even between self-hosting and an API works through where self-hosting starts to win. For a product team whose inference sits in the critical path, the serverless route usually keeps the switch reversible, which is the property the next section protects.

Staging the switch so it stays reversible

A switch you cannot undo is a bet. A switch you can undo is an experiment, and the difference is a written rollback condition and a staged rollout. Google's SRE workbook defines canarying as a partial and time-limited deployment of a change in a service, whose evaluation helps decide whether or not to proceed with the rollout. Apply that structure to the provider, not just to your own code:

  1. Write measurable acceptance and rollback conditions before moving traffic.
  2. Use permitted shadow or replay data, discard candidate responses, and disable duplicate tool actions.
  3. Route a small share of user traffic only after the candidate passes the agreed checks.
  4. Increase traffic while monitoring quality, errors, latency and cost.
  5. Retire the old route only after the agreed rollback window closes.

Mirroring production prompts sends data to a second processor. Check its retention, access and processing terms first. Lyceum self-asserts zero retention for inference and lists hosting per model. The model’s EU label does not cover your own gateway logs, external tools or other parts of the application.

Lyceum’s serverless API provides pre-hosted models billed per token without a GPU reservation. Use it as a candidate in the same evaluation process as any other provider. Estimate all 5 cost components, test the actual endpoint and keep rollback possible.