Claude Code speaks the Anthropic API
Yes, Claude Code runs directly against your own custom model endpoints. Anthropic's model-configuration documentation confirms that the client natively supports routing requests to a custom base URL. The route takes just three commands, with no proxy required.
- Native protocol compatibility: Claude Code communicates directly with the endpoint using standard Anthropic request headers and message structures.
- Subagents included: with lyceum code, the subagents Claude Code spawns run on the same Lyceum model as the main session.
What you need before you start
Before launching Claude Code against a custom endpoint, ensure your local development environment meets the base requirements for the CLI tooling and authentication.
- Python 3.10 or newer: Required to run the CLI helper package across Linux, macOS, or WSL environments.
- An account and API key: Generate an API key by logging into the dashboard, navigating to Account, and selecting API keys. The secret key is displayed only once upon generation, so store it securely.
- Claude Code installed: Ensure Claude Code is installed on your workstation via the official native installer or package managers.
For AI consultancies and implementation partners, one setup that every engineer can reproduce gives the team a repeatable way to run Claude Code on open-weight models across client engagements, without changing individual project codebases.
Install the Lyceum CLI and launch Claude Code on Kimi K3
The recommended route is the Lyceum CLI. It keeps its own Claude Code profile and launches Claude Code with the Lyceum endpoint and your API key already set.
Run the following three commands in your terminal:
pip install lyceum-cli
lyceum auth login --api-key lk_your_api_key_here
lyceum codeExecuting lyceum code starts Claude Code in an isolated profile directory (~/.lyceum/claude) that operates independently of any existing Anthropic credentials. By default, the session targets moonshotai/kimi-k3.
When Claude Code launches for the first time, it prompts in the terminal whether to approve the configured API key. Select Yes to proceed. If you need to target a different model for a specific session, pass the --model flag, or append --save to persist the selection as your new profile default:
lyceum code --model z-ai/glm-5.2 --saveOn Serverless Inference, Kimi K3 is billed on transparent per-token consumption at $3.00 per 1M input tokens, $0.75 per 1M cached input tokens, and $15.00 per 1M output tokens, with no minimum spend or base platform subscription fees. For full parameter benchmarks and context length breakdowns, review our guide on the Kimi K3 API: where to run it and what a 1M-token context costs. You can also reference the two-minute Lyceum setup video for a complete terminal walkthrough.
| Command Option | Action | Scope |
|---|---|---|
| lyceum code | Launches Claude Code on moonshotai/kimi-k3 | Main session and subagents |
| --model <id> | Overrides the active model for the current execution | Active session |
| --save | Persists the selected model to the local profile configuration | Future CLI sessions |
The two environment variables, without the CLI
If you prefer not to install the CLI, or if you are configuring automated CI/CD runners, Docker containers, and background scripts, you can configure Claude Code directly using two standard environment variables:
export ANTHROPIC_BASE_URL="https://api.lyceum.technology/anthropic"
export ANTHROPIC_AUTH_TOKEN="lk_your_api_key_here"
claudeIt is critical to export ANTHROPIC_AUTH_TOKEN rather than ANTHROPIC_API_KEY. When ANTHROPIC_API_KEY is detected, Claude Code initiates an interactive security prompt requiring terminal user input to confirm the custom key on initial startup. In non-interactive environments, such as headless CI runners or piped execution with claude -p, this interactive prompt cannot be answered. The execution terminates immediately with the error: Not logged in · Please run /login.
Setting ANTHROPIC_AUTH_TOKEN bypasses the interactive approval prompt entirely, allowing Claude Code to authenticate directly against the gateway in automated pipelines.
| Environment Variable | Interactive Prompt Triggered | CI / Headless Compatibility |
|---|---|---|
| ANTHROPIC_AUTH_TOKEN | No (Direct authentication) | Compatible with non-interactive scripts and CI |
| ANTHROPIC_API_KEY | Yes (Requires manual terminal confirmation) | Fails with 'Not logged in · Please run /login' |
Switching models mid-session and the tier mapping
When using the direct environment variable route, Claude Code requests models using standard Anthropic tier aliases such as opus, sonnet, and haiku. The gateway translates these tier requests automatically to open-weight model architectures.
You can override tier defaults at any time during an active session by issuing the /model slash command followed by an explicit model string, without restarting your terminal or losing conversation context.
| Claude Tier / Alias | Default Model | Exact Model ID | Recommended Workload |
|---|---|---|---|
| opus / fable | GLM-5.2 | z-ai/glm-5.2 | Complex multi-file refactoring and architectural planning |
| sonnet | Kimi-K2.7-Code | moonshotai/kimi-k2.7-code | Daily agentic software engineering and code generation |
| haiku | MiniMax-M3 | minimax/minimax-m3 | Fast repository exploration, linting, and unit testing |
| custom override | Kimi-K3 | moonshotai/kimi-k3 | Long-context reasoning and large codebase ingestion |
To switch models dynamically during a session, enter the command directly into the Claude Code prompt:
/model moonshotai/kimi-k3
/model moonshotai/kimi-k2.7-code
/model z-ai/glm-5.2Note two operational rules when switching models. First, do not append Anthropic's [1m] context suffix (for example, sonnet[1m]); Claude Code will not connect through the Lyceum endpoint with it set. Use sonnet instead. Second, when using the lyceum code CLI wrapper, issuing /model changes the model for the main interactive session only. To update child subagents as well, exit the session and restart with lyceum code --model <id>.
For deeper comparisons across open architectures, see our guide on the best open model APIs for agentic coding, or review our walkthrough on Using Lyceum Models in opencode.
Where these setups usually break
When connecting Claude Code to custom endpoints, infrastructure engineers typically encounter two specific failure modes that make a functioning endpoint appear broken.
The reasoning token budget trap
Open reasoning models return intermediate chain-of-thought tokens in a distinct reasoning field or output channel. If a client configuration sets a low max_tokens parameter, the model can spend its entire token allocation on internal reasoning before emitting any final code text. In Claude Code, this manifests as an empty response or an abrupt completion. To resolve this, use the -instant model variants or leave enough output token headroom. For a technical breakdown of this mechanism, read our guide on why a reasoning model returns an empty response.
Stored rejection at the API key prompt
When launching lyceum code or running Claude Code with a custom key for the first time, selecting No at the terminal approval prompt stores a rejection state in the local profile cache. Claude Code stores that choice and never asks again, so requests no longer reach Lyceum. Because this choice is cached in the user directory, restarting the command will not clear the state.
To clear the cached rejection and reset the profile, purge the local Claude directory and relaunch:
rm -rf ~/.lyceum/claude
lyceum code| Failure Mode | Observed Symptom | Resolution |
|---|---|---|
| Reasoning token exhaustion | Empty response text or truncated output | Switch to instant model variants or expand token limit |
| Cached prompt rejection | Requests never reach Lyceum | Run rm -rf ~/.lyceum/claude and relaunch lyceum code, selecting 'Yes' |
Confirm the model and verify the endpoint
To verify that your installation is actively routing to the custom endpoint rather than default cloud endpoints, inspect the terminal launch banner when executing lyceum code. The banner explicitly prints the resolved model string and says it applies to the main session and subagents.
You can also verify gateway connectivity and authentication outside of Claude Code using a single curl command against the Messages API endpoint:
curl https://api.lyceum.technology/anthropic/v1/messages \
-H "x-api-key: lk_your_api_key_here" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "moonshotai/kimi-k3",
"max_tokens": 128,
"messages": [{"role": "user", "content": "Respond with: Lyceum endpoint verified."}]
}'A successful request returns a standard Anthropic message JSON payload containing the model identifier, token usage statistics and the response content. Reasoning models such as Kimi K3 return a thinking block before the text block (trimmed here):
{
"id": "msg_01AbCdEfGhIjKlMnOpQrStUv",
"type": "message",
"role": "assistant",
"model": "moonshotai/kimi-k3",
"content": [
{
"type": "thinking",
"thinking": "The user wants me to respond with exactly \"Lyceum endpoint verified.\"",
"signature": ""
},
{
"type": "text",
"text": "Lyceum endpoint verified."
}
],
"stop_reason": "end_turn",
"usage": {
"input_tokens": 18,
"output_tokens": 7
},
"stop_sequence": null
}Serverless Inference provides on-demand access to open-weight models through OpenAI and Anthropic compatible endpoints with per-token billing and zero base fees.
Create an API key in the Lyceum dashboard, run lyceum code in your own repo, and watch the two-minute Lyceum setup video if you want to see it end to end first.