What opencode is, and why a custom endpoint
Terminal-based AI coding agents have become standard tools in modern engineering workflows. Among these tools, opencode provides a fast, keyboard-driven interface that executes directly inside your development shell, analyzing multi-file repositories, applying diffs, and orchestrating terminal tasks without leaving the command line. Under the hood, opencode relies on the AI SDK and Models.dev to interface with over 75 language model providers, giving developers the architectural freedom to route generation requests away from default proprietary vendors and toward custom inference infrastructure.
The architectural case for open model endpoints
For engineering organizations operating at scale, routing coding agent traffic to private or sovereign infrastructure resolves two persistent operational hurdles: escalating token costs and cross-border data transfer risks. High-frequency agentic loops generate massive token volumes through repeated workspace indexing, system prompt retransmission, and tool-call retries. Standard commercial coding assistants charge premium rates for these cycles, creating significant operational expenditure that compounds as engineering teams expand.
By pointing opencode at custom OpenAI-compatible endpoints, teams can substitute closed models with frontier open-weight models that deliver comparable reasoning at a fraction of the cost per million tokens. While our earlier coverage explored configuring alternatives such as the Claude Code setup guide, setting up opencode provides a completely open-source terminal client that natively supports drop-in OpenAI-compatible provider definitions.
Beyond token economics, custom endpoint configuration allows European engineering teams to enforce strict data governance. By standardizing on infrastructure hosted within European jurisdictions, organizations avoid unvetted data egress, comply with internal security mandates, and prevent proprietary source code from being retained for third-party model retraining.
The provider block in opencode.json
Configuring a custom endpoint in opencode requires defining a dedicated provider entry inside your opencode configuration file. opencode lets you customise providers through the provider section of its config, including the baseURL used for proxy services or custom endpoints. To connect to an OpenAI-compatible serverless endpoint, you attach the npm package @ai-sdk/openai-compatible, which the AI SDK documents as the package you use for language model providers that implement the OpenAI API. Its apiKey setting, when specified, adds an Authorization header to request headers with the value Bearer followed by the key.
Configuration schema and endpoint mapping
The opencode config docs put global, user-wide configuration in ~/.config/opencode/opencode.json and per-project configuration in an opencode.json file in your project root, and they state that config files are merged rather than replaced, with project config taking the highest precedence among the standard config files. Inside the provider block you point the baseURL at your OpenAI-compatible endpoint. Authentication is handled via a standard Bearer token mapped directly from the local shell environment.
The following JSON configuration demonstrates how to declare the custom provider, attach the required npm driver, specify the base URL, and define the environment variable mapping for the authentication key:
{ "$schema": "https://opencode.ai/config.json", "provider": { "eu-inference": { "npm": "@ai-sdk/openai-compatible", "options": { "baseURL": "https://api.lyceum.technology/openai/v1", "apiKey": "{env:LYCEUM_API_KEY}" }, "models": { "deepseek/deepseek-v4-pro-0813": { "name": "DeepSeek V4 Pro", "limit": { "context": 200000, "output": 65536 } } } } } }
Within this structure, the provider key establishes the identifier namespace used across the CLI. Any model configured under this block is referenced using the provider/model-id syntax, ensuring clean separation from default cloud providers.
Export the key:.env files are not read
The most frequent setup stumbling block encountered by developers configuring opencode is authentication resolution. Developers frequently place their API credentials inside a local.env file in the root of their workspace, expecting the CLI to parse the file on startup. In practice, the {env:LYCEUM_API_KEY} placeholder is resolved against the environment of the process that launches opencode, so a key that only exists in a project.env file is never seen and the request goes out without a usable bearer token.
Setting persistent shell environment variables
When opencode parses the provider configuration string {env:LYCEUM_API_KEY}, it searches the active process environment. If the variable is absent, requests to the endpoint fail immediately with an authentication error or an empty bearer token header. To ensure that the key is consistently available across all terminal sessions, the export must be defined in your interactive shell configuration file.
- Open your primary shell configuration file in a text editor, typically ~/.zshrc for Zsh or ~/.bashrc for Bash.
- Add the export statement: export LYCEUM_API_KEY="lk_your_api_key_here"
- Save the configuration file and reload your shell environment with source ~/.zshrc or source ~/.bashrc.
- Verify that the variable is active in your current session by running printenv LYCEUM_API_KEY.
Setting this variable globally ensures that both manual CLI commands and background agent subshells inherit the correct authentication headers without manual re-exporting during active development cycles.
Set context limits or compaction miscounts
When querying upstream model discovery endpoints, compatibility variations between inference engines can impact client-side context management. Standard OpenAI-compatible /models endpoints return lists of available model identifiers but omit metadata regarding exact token context windows and output sequence boundaries. opencode's models documentation shows per-model settings being configured as keys under a provider's models section in the config file, so leaving that metadata out means the client works from its own assumptions rather than the model's real window.
Preventing compaction errors with explicit limits
If opencode assumes a much smaller token limit than the model actually supports, it will trigger aggressive context compaction far too early, compressing or discarding valuable repository context, file diffs, and previous conversational history. Conversely, if it overestimates capacity on a smaller model, long agentic loops will fail with server-side bad request exceptions when the context boundary is crossed.
To guarantee stable memory management across long sessions, you must supply explicit limit parameters in each model definition inside opencode.json. These numbers should be taken directly from the official model specifications rather than inferred from the API. Understanding these architectural nuances is critical when adapting client configurations, as detailed in our analysis of OpenAI Compatible APIs: What Breaks When Switching Models.
| Model Identifier | Target Workload | Context Limit (Tokens) | Output Limit (Tokens) |
|---|---|---|---|
| deepseek/deepseek-v4-pro-0813 | Complex Refactoring & Architecture | 200000 | 65536 |
| moonshotai/kimi-k2.7-code | High-Throughput Agentic Coding | 200000 | 65536 |
| moonshotai/kimi-k3 | Multi-File Reasoning & Synthesis | 200000 | 65536 |
| z-ai/glm-5.2 | General Scripting & Unit Tests | 200000 | 65536 |
| deepseek/deepseek-v4-flash-0731 | Fast Edits & Commit Generation | 128000 | 32768 |
By declaring context and output ceilings explicitly in the configuration block, opencode maintains accurate rolling token calculations, ensuring prompt pruning and compaction execute exactly at designated buffer thresholds.
Choosing a model for agentic coding
Selecting the appropriate model configuration for opencode depends on the specific engineering task, code complexity, and latency requirements. In agentic development loops, the language model does more than complete text: it inspects directory trees, parses compiler outputs, constructs patch diffs, and calls local tool interfaces. That matters because, as opencode's own documentation puts it, there are only a few models that are good both at generating code and at calling tools.
Matching model architecture to agentic tasks
For complex multi-file refactoring, dependency migrations, and deep architectural design, frontier reasoning architectures such as deepseek/deepseek-v4-pro-0813, moonshotai/kimi-k3, and z-ai/glm-5.2 provide high tool-call precision and structured JSON compliance. These models accurately interpret repository layouts, maintain state across deep call stacks, and produce cleanly formatted git patches without hallucinations.
For high-speed, day-to-day code completion, single-function implementation, and automated test generation, optimized coding variants like moonshotai/kimi-k2.7-code and deepseek/deepseek-v4-flash-0731 deliver rapid time-to-first-token and lower operational cost while clearing strict quality benchmarks. Engineering teams can evaluate these trade-offs systematically by referencing our methodology in Finding the Cheapest Open Model That Clears Your Quality Bar.
When configuring model identifiers in opencode.json, ensure that each model string matches the active API roster exactly, using the vendor-prefixed format. Model hosting regions and data residency commitments remain model-specific properties; teams requiring strict European residency should verify the hosting characteristics of each individual model configuration on its respective specification page.
Strict JSON and the default-model surprise
Two additional operational nuances in opencode require careful attention during configuration: JSON parser strictness and default model selection behavior. Overlooking either of these mechanics can lead to a config that never loads or to requests being routed at a provider you did not intend.
JSON schema validation and JSONC support
Standard opencode.json configuration files are parsed using strict JSON specifications. This means that trailing commas after array items or object keys, as well as JavaScript-style single-line or block comments, will cause the parser to fail on initialization. When opencode encounters invalid JSON syntax, it aborts loading the custom configuration and reverts to base system defaults without detailed terminal error logging.
- If your team prefers annotating configuration files with inline comments or maintaining trailing commas, rename the configuration file from opencode.json to opencode.jsonc.
- Validate that all brackets and quotation marks are properly balanced before restarting the CLI.
- Verify that environment variable interpolation strings use the exact {env:VAR_NAME} syntax without extra whitespace.
Enforcing the top-level model selection
Another frequent misconfiguration involves omitting the top-level model key in the configuration root. When opencode starts, it looks for models in this priority order: the --model or -m command line flag, then the model set in the opencode config, then the last used model, then the first model using an internal priority. Without an explicit model key, that last step can quietly route your queries to a provider other than your custom endpoint.
To prevent this fallback behavior, explicitly define the default model at the root level of your opencode.json or opencode.jsonc file:
{ "$schema": "https://opencode.ai/config.json", "model": "eu-inference/deepseek/deepseek-v4-pro-0813" }
Verify the connection with opencode models
Once your configuration file is saved and your environment variables are active, verifying the connection takes only a few terminal commands. The opencode CLI reference documents a run command that takes a prompt directly, a --model / -m flag described as the model to use in the form provider/model, and a models command that lists all available models from configured providers and displays them in provider/model format.
Listing and testing configured endpoints
To confirm that your custom provider block has loaded and that model identifiers have registered correctly, run the models listing command in your terminal:
opencode models
The terminal output should display your configured models under the custom provider prefix, resolving as provider-id/deepseek/deepseek-v4-pro-0813 or whichever model strings you declared under that provider key. If the models do not appear, check that your configuration file resides in the correct directory path and contains valid JSON syntax.
Next, execute a one-shot inference test to validate network connectivity, authentication headers, and response streaming from the command line:
opencode run -m eu-inference/deepseek/deepseek-v4-pro-0813 "Reply with one word: pong"
A successful response confirms that the request traversed the custom endpoint, validated the bearer token, executed token generation, and streamed the output back to your active shell session.
Production rollout with Serverless Inference
Routing opencode through Lyceum Serverless Inference provides engineering organizations with a direct path to lower developer infrastructure spend while maintaining absolute sovereignty over codebases and runtime data. By running pre-hosted open-weight frontier models behind an OpenAI-compatible API, teams eliminate dedicated GPU provisioning overhead, avoid idle infrastructure costs, and ensure all prompt traffic processes under European data governance standards. For enterprise teams standardizing their agentic workflows, Serverless Inference delivers transparent per-token billing, sub-second time-to-first-token, and full architectural flexibility across the modern developer stack. A short Lyceum walkthrough shows the whole setup end to end: creating an API key, installing the Lyceum CLI and starting Claude Code on Kimi K3 in about two minutes.