Copilot now accepts a custom endpoint
Enterprise development teams scaling artificial intelligence tooling across hundreds of engineering seats face a steep cost curve with fixed per-seat subscriptions. GitHub Copilot addresses this cost pressure by adding Bring Your Own Key (BYOK) support through a Custom Endpoint provider. This mechanism replaces the legacy OpenAI Compatible configuration, giving organisations direct control over model selection, per-token expenditure, and backend compute infrastructure.
For cost-conscious enterprise adopters, BYOK allows switching the chat interface from proprietary closed-source models to open-weight models without disrupting developer workflows. Connecting your own provider lets engineering teams maintain their familiar Visual Studio Code interface while directing token consumption to cost-effective endpoints. Teams exploring open-source coding backends often evaluate the best open model APIs for agentic coding to align token economics with developer throughput.
Custom endpoints require the correct model identifier, endpoint path and token limits. Our guide on Using Lyceum Models in opencode: Custom Provider Setup covers another client. For Copilot, use the slash-free alias below and check request errors as well as empty replies when troubleshooting.
- Replaces the deprecated github.copilot.chat.customOAIModels configuration with unified Custom Endpoint architecture.
- Bills external model usage through your provider; it does not cancel an existing Copilot subscription.
- Works in VS Code chat; Agent Host sessions have a separate experimental switch.
- Requires a supported model identifier and suitable output limits.
Adding the endpoint in VS Code
Connecting VS Code Copilot to an external endpoint takes place through the Language Models editor overlay. You access this management interface through the command palette or directly within the Chat view. Setting up the provider registers your endpoint across workspace sessions, allowing developers to switch between internal foundation models and standard tooling.
To open the configuration interface, press Ctrl+Shift+P (or Cmd+Shift+P on macOS) and execute the Chat: Manage Language Models command. Alternatively, click the gear icon located inside the language model picker at the base of the GitHub Copilot Chat panel. From the modal overlay, click Add Models and select Custom Endpoint.
Enter the provider details, then configure the model in the file VS Code opens:
- Group Name: Enter a descriptive label for clear grouping in the UI.
- Display Name: Provide an identifiable model tag, such as GLM 5.2 Instant.
- API Key: Enter your platform key or supply an input variable token like ${input:lyceumApiKey}.
- API Type: Select Chat Completions from the dropdown list.
- Model URL: In chatLanguageModels.json, set url to https://api.lyceum.technology/openai/v1/chat/completions.
Use the following Chat Completions configuration as a starting point. These conservative input and output limits are client settings, not a claim about the model’s maximum context. Increase them only within the current model and service limits.
[
{
"name": "Custom Provider",
"vendor": "customendpoint",
"apiKey": "${input:lyceumApiKey}",
"apiType": "chat-completions",
"models": [
{
"id": "z-ai-glm-5.2-instant",
"name": "GLM-5.2 Instant",
"url": "https://api.lyceum.technology/openai/v1/chat/completions",
"toolCalling": true,
"vision": false,
"maxInputTokens": 100000,
"maxOutputTokens": 8192
}
]
}
]After saving chatLanguageModels.json, reload the window or restart VS Code to ensure the language model picker populates the newly registered models.
Why the model ID needs no slash
The most common stumbling block when configuring custom providers in Copilot is model identifier formatting. Standard open-source model registries, including Hugging Face and vLLM inference deployments, follow the organisation/model-name convention (such as z-ai/glm-5.2 or moonshotai/kimi-k2.7-code). However, the VS Code Copilot client rejects model identifiers containing a forward slash (/) when loading custom endpoints.
Lyceum documents slash-free aliases for clients that reject forward slashes. Use z-ai-glm-5.2-instant here. This alias mapping is a Lyceum feature; do not assume that replacing slashes works with another provider.
When mapping your models, keep the id field slash-free while assigning whatever human-readable string you prefer to the name field (for instance, mapping z-ai/glm-5.2 to z-ai-glm-5.2). Teams standardising multi-model setups must understand what breaks when you switch models on an OpenAI-compatible API for deeper insight into request schema translations.
Why to pick an instant variant
When developers connect advanced open reasoning models like GLM-5.2 directly to Copilot Chat, they frequently experience the empty response trap: the model appears to process the request for several seconds, but the UI returns a blank response bubble. This issue stems from architectural friction between native reasoning token streams and Copilot's token processing.
Reasoning-enabled backends can return a separate reasoning field before the visible answer. Rendering and parameter support differ between clients and versions. Current VS Code supports modelOptions, so a blanket claim that custom requests cannot include reasoning controls is outdated.
If reasoning exhausts the generation budget, the response can end before visible text appears. That is one possible cause of an empty reply. Check the finish reason, request errors and selected model before changing the token limit.
- Reasoning and visible answers can share an output budget.
- For this guide, use z-ai-glm-5.2-instant to avoid a separate reasoning phase.
- An instant suffix is not a universal guarantee: verify the behaviour of the specific variant.
- Keep enough output capacity for the requested code or explanation.
Use the GLM-5.2 Instant alias shown here rather than adding -instant to an arbitrary model ID. Check the live model list before changing models.
What bring-your-own-key does not cover
BYOK changes who serves and bills the selected model. It does not automatically replace every Copilot feature or an organisation’s existing subscription.
| Copilot Capability | BYOK Custom Endpoint Support | Operational Requirements & Caveats |
|---|---|---|
| Copilot Chat Panel | Supported | Fully supported via Chat Completions or Messages API |
| Commit Message Generation | Supported | Requires configuring chat.utilitySmallModel to the BYOK model |
| PR Title & Description | Supported | Set chat.utilitySmallModel |
| Agent use in chat | Supported | Requires a tool-capable model |
| Agent Host sessions | Experimental | Also enable chat.agentHost.byokModels.enabled |
| Inline Code Completions | Not Supported | Ghost text completions remain strictly on GitHub native models |
| Semantic Workspace Search | Not Supported | Separate GitHub account and feature access requirements apply |
Agent use requires tool calling. The experimental chat.agentHost.byokModels.enabled setting applies specifically to Agent Host sessions. Business or Enterprise administrators can disable BYOK through policy.
Inline code completions, semantic search and embedding-based features are outside this custom model setup. Check their separate account and plan requirements before changing your team’s subscriptions.
In addition to the standard Chat Completions interface, VS Code Copilot's Custom Endpoint provider supports the Anthropic Messages API protocol. This alternative transport is useful for organisations maintaining unified Anthropic SDK client configurations across their internal developer tooling.
For Messages, set apiType to messages and use the full URL https://api.lyceum.technology/anthropic/v1/messages. Keep z-ai-glm-5.2-instant as the model ID rather than relying on a Claude tier alias. The corresponding Claude Code setup is covered in the linked documentation.
[
{
"name": "Anthropic Provider",
"vendor": "customendpoint",
"apiKey": "${input:lyceumApiKey}",
"apiType": "messages",
"models": [
{
"id": "z-ai-glm-5.2-instant",
"name": "GLM 5.2 Instant (Messages API)",
"url": "https://api.lyceum.technology/anthropic/v1/messages",
"toolCalling": true,
"vision": false,
"maxInputTokens": 100000,
"maxOutputTokens": 8192
}
]
}
]Messages requests use x-api-key authentication by default; Chat Completions uses a Bearer token. Both routes accepted the alias and returned a short answer in API checks on 21 September 2026. This confirms endpoint access, not a full Copilot agent workflow.
Compare external token spend with the Copilot features your team still needs before changing licences.
Serverless Inference offers managed access to open-weight models with per-token billing. OpenAI-compatible requests and the Anthropic Messages bridge provide alternative interfaces. Check the chosen model’s support for tools, structured output and images before relying on those features.
- Confirm serving regions and any fallback routes with Lyceum before making a data residency commitment.
- Validate streaming, tool calls and structured output for the chosen model and client.
- Check the current price book, including any cached-input tier.
- Check status.lyceum.technology for service status and agree any contractual availability requirements with sales.
To connect your engineering team's VS Code Copilot instances to Lyceum, generate an API key in the Lyceum dashboard and add your endpoint configuration to get started.