Zed declares model providers in settings.json
Zed supports provider configuration in settings.json and through Agent Settings. Use agent: open settings to add a provider in the LLM Providers section. Keep credentials separate from shareable model configuration.
This guide shows how to add an open-weight Lyceum model to Zed's agent panel with a single settings.json block. When evaluating open architectures, understanding the best open model APIs for agentic coding helps teams balance per-token execution economics against reasoning depth.
The configuration options declared under language_models configure the native Zed Agent and editor-owned AI features such as inline transformations and thread summaries. They do not configure External Agents or Terminal Threads, which manage their own runtime authentication and connection configurations independently.
- Centralized configuration: Model settings sit in standard JSON that can be versioned and reviewed.
- Native agent integration: Powers the built-in Zed Agent panel, inline generation, and commit message synthesis.
- Scattered runtime separation: Keeps the editor's core model definitions separate from standalone external CLI agents.
Adding the Lyceum provider block
Open Zed’s settings file from the command palette and merge the block below into your existing language_models object. Alternatively use agent: open settings, then Add Provider in LLM Providers. Do not replace unrelated settings.
Add the custom provider definition under the language_models.openai_compatible key. The api_url property points to the standard endpoint https://api.lyceum.technology/openai/v1. Each entry in the available_models array defines the model identifier, human-readable display label, context sequence limit, and maximum output token allocation.
{
"language_models": {
"openai_compatible": {
"lyceum": {
"api_url": "https://api.lyceum.technology/openai/v1",
"available_models": [
{
"name": "z-ai/glm-5.2",
"display_name": "GLM-5.2 (Lyceum)",
"max_tokens": 1000000,
"max_output_tokens": 8192,
"reasoning_effort": "high",
"capabilities": {
"tools": true,
"images": false,
"parallel_tool_calls": false,
"prompt_cache_key": false,
"chat_completions": true,
"interleaved_reasoning": true,
"max_tokens_parameter": true
}
}
]
}
}
}
}When transitioning from other client environments, such as Using Lyceum Models in opencode: Custom Provider Setup, defining explicit model identifiers ensures the editor resolves the target endpoint without silent fallbacks.
Declaring limits and capabilities
Zed supplies defaults for custom provider capabilities. Explicit settings make this example easier to review. Keep chat_completions enabled for Lyceum’s documented chat endpoint, and match tool, image and reasoning settings to the selected model.
The capabilities configuration maps directly to specific inference capabilities:
- tools: keep true for native agent tool calls, then test a real read/edit cycle
- images: use false for GLM-5.2; enable only for a model and endpoint that accept images
- parallel_tool_calls: false is a conservative starting setting, not a universal limitation of open models
- prompt_cache_key: leave false unless the endpoint documents support for that request parameter
The dashboard checked on 1 October 2026 lists the context and input types below. Lyceum documents a maximum output request allowance of 65,536 tokens, including reasoning. That is an upper limit, not a required setting. The example requests 8,192; adjust it within the endpoint limit and your budget.
| Model identifier | Advertised context | Output request ceiling | Advertised input |
|---|---|---|---|
| z-ai/glm-5.2 | 1,000,000 | 65,536 | Text |
| z-ai/glm-5.3 | 1,000,000 | 65,536 | Text |
| minimax/minimax-m3 | 1,000,000 | 65,536 | Text and images |
| qwen/qwen3.5-9b | 256,000 | 65,536 | Text and images |
| moonshotai/kimi-k2.7-code | 256,000 | 65,536 | Text, images and PDF |
Keeping the key out of version control
Never commit authentication tokens or secrets inside settings.json. Because settings.json is often tracked in version control or distributed across public dotfiles repositories, embedding an API token directly in this file creates an immediate security exposure.
Zed decouples provider declaration from secret storage. When you open the agent panel or trigger a completion with a configured custom provider for the first time, Zed prompts you for the API key in the UI and stores it securely in your operating system keychain (such as macOS Keychain or Linux Secret Service).
- Add the provider configuration without a credential field
- Open agent: open settings and select the custom Lyceum provider
- Enter your Lyceum API key through the provider UI
- Select the model and test a short message before trying a file-edit task
For provider ID lyceum, Zed reads LYCEUM_API_KEY from its local process environment. Non-empty environment values take precedence over saved keys. Make sure the Zed process receives the variable; changing a terminal variable does not necessarily change an already-running GUI app.
When the agent loop stalls
A stalled agent or HTTP 400 needs the actual error and request context. Check authentication, the model identifier, request limits and tool schema. Temporarily disabling parallel tool calls can help isolate a compatibility issue, but is not a diagnosis by itself.
OpenAI-compatible endpoints differ in feature behaviour. Test an ordinary tool call, its returned result and the next model turn. Lyceum’s documentation also notes model-specific forced tool-choice limits. Do not infer an absence of tool support from one forced-call failure.
| Symptom | Check | Next step |
|---|---|---|
| HTTP 401 | Key, account and environment precedence | Use the intended active key |
| HTTP 400 | Exact error, model, schema and token limits | Correct the rejected field before retrying |
| Prose instead of a tool call | Tool configuration and model response | Test a simple tool with normal automatic selection |
| Unexpected truncation | Context, output cap and reasoning usage | Adjust within documented limits and inspect compaction |
Handling reasoning models and output fields
Modern open architectures frequently incorporate reasoning phases where the model generates a thinking trace before producing its final answer text. When querying reasoning models on an OpenAI-compatible surface, the runtime returns reasoning content in a separate field (such as reasoning_content) distinct from the main message content field.
Zed handles OpenAI-compatible chat completions by parsing standard message fields. Because reasoning tokens count toward total generated sequence limits, you must ensure that max_output_tokens is set large enough to accommodate both the reasoning trajectory and the resulting code generation.
- Reasoning: verify the response field and Zed’s interleaved_reasoning setting for the chosen model
- Output: choose a practical cap at or below 65,536, remembering that reasoning consumes part of it
- Variants: use only currently listed identifiers and documented reasoning controls; do not invent an instant suffix
Serverless Inference for enterprise teams
Migrating developer tooling to open-weight models allows engineering organizations to replace expensive per-seat licenses with usage-based infrastructure. By standardizing Zed configurations across team repositories, companies can provide high-capability coding agents to every developer without managing complex client setups or operating dedicated GPU clusters.
Lyceum Serverless Inference provides pre-hosted open models accessible through a standard OpenAI-compatible interface with per-token pricing and no base fees. For organizations evaluating availability, Lyceum publishes no SLA for serverless inference; real-time operational status is tracked at status.lyceum.technology.
- Usage pricing: account for input, cached input and output, including billable reasoning
- Shared configuration: distribute the tested settings while each developer supplies a key
- Compatibility: verify the agent’s actual tool and streaming workflow before a team-wide rollout
Get an API key in the Lyceum dashboard and try it on your own repository.