Z.ai
GLM-5.3
Z.ai’s large model for complex coding and long agent tasks, built on the same base model as GLM-5.2.
- Context window
- 1M tokens
- Accepts
- Text
- Input per 1M tokens
- $1.40
- Output per 1M tokens
- $4.40
Send your first request
Set LYCEUM_API_KEY to your API key, then run this request. Usage is billed to your account.
- Base URL
https://api.lyceum.technology/openai/v1- Model ID
z-ai/glm-5.3- Request fields
model, messages, stream, max_tokens, tools, tool_choice
Documented request fields for this endpoint. See the documentation for model-specific controls and limits.
app.py · Install openai and set LYCEUM_API_KEY
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["LYCEUM_API_KEY"],
base_url="https://api.lyceum.technology/openai/v1",
)
response = client.chat.completions.create(
model="z-ai/glm-5.3",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)app.ts · Install openai and set LYCEUM_API_KEY in Node.js
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.LYCEUM_API_KEY,
baseURL: "https://api.lyceum.technology/openai/v1",
});
const response = await client.chat.completions.create({
model: "z-ai/glm-5.3",
messages: [{ role: "user", content: "Hello!" }],
});
console.log(response.choices[0]?.message.content);Terminal · Set LYCEUM_API_KEY in your terminal
curl https://api.lyceum.technology/openai/v1/chat/completions \
-H "Authorization: Bearer $LYCEUM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "z-ai/glm-5.3",
"messages": [
{"role": "user", "content": "Hello!"}
]
}'What it costs and how fast it runs
You pay only for the tokens you use. Speed figures come from our own status checks.
Token pricing
| USD per million tokens | |
|---|---|
| Input | $1.40 |
| Cached input | $0.26 |
| Output | $4.40 |
No base fee. Caching applies only where listed. Check the model's processing region before sending data.
Speed on Lyceum
| Last 7 days | |
|---|---|
| Time to first token | 294 ms |
| Output speed | 80 tokens/s |
Median over the last 7 days from our status checks: short requests. Measured 6 October 2026, 22:38 UTC. Reasoning tokens count towards both figures. Live status and availability
The model in detail
Checked against the sources below. Anything they don't state is left out.
Capabilities
- Input and output
- Text in, text out
- Context window
- 1M tokens
- Reasoning
- On by default
- Tool calling
- Supported
Architecture and licence
- Total parameters
- 744B
- Active parameters
- 40B
- Architecture
- Mixture of experts
- Licence
- GLM-5.3 License
- Released
- 14 August 2026
Availability
- API model ID
z-ai/glm-5.3- Provider
- Z.ai
- Processing region
- EU-hosted
- Input and output types: official model card, checked 10 September 2026.
- Capabilities and request fields: Lyceum API documentation, checked 9 September 2026.
- Parameters, licence and release date: Z.ai’s own sources, including its model card and announcement, checked 29 September 2026.
Common questions about GLM-5.3
Short answers about the model ID, price, limits and behaviour.
What is the model ID for GLM-5.3?
Use z-ai/glm-5.3 as the model value. The OpenAI-compatible base URL is https://api.lyceum.technology/openai/v1. You authenticate with your Lyceum API key as a Bearer token.
How much does GLM-5.3 cost?
GLM-5.3 costs $1.40 per 1M input tokens, $0.26 per 1M cached input tokens and $4.40 per 1M output tokens. Prices are in US dollars. Billed per token. No base fee.
What is the context window of GLM-5.3?
GLM-5.3 has a 1M tokens context window. One request can generate up to 65,536 tokens and run for up to 300 seconds. Set stream: true for long outputs.
Can GLM-5.3 read images?
No, not today. Image requests are rejected with a 400 error. Pick a model that reads images if you need that.
Can I turn off reasoning for GLM-5.3?
Not cleanly today. Reasoning is on by default, and switching it off moves the reasoning into the answer text instead of removing it. Leave it on and set max_tokens high enough for both the reasoning and the answer.
Does GLM-5.3 support tool calling?
Yes. Send tools in the standard OpenAI format. Automatic tool calls, parallel calls and tool-result round trips work. It does not enforce tool_choice: "required" or a named function. It can answer in plain text with no tool calls, so check message.tool_calls and fall back when it is empty.
Does GLM-5.3 support prompt caching?
Yes, and it is automatic. When requests share an identical start, the cached part is billed at $0.26 per 1M tokens instead of $1.40. Keep the stable part of your prompt at the front. A cache hit is best effort, so read usage.prompt_tokens_details.cached_tokens, which is missing on a miss.
Where does GLM-5.3 run, and is my data kept?
GLM-5.3 runs on EU-hosted infrastructure. Prompts and outputs are processed, not stored, and never used for training.
How do I call GLM-5.3 with the OpenAI SDK?
Point the OpenAI SDK at https://api.lyceum.technology/openai/v1, use your Lyceum API key and set model="z-ai/glm-5.3". Chat completions, streaming, tool calling and usage accounting work as they do with OpenAI.
Behaviour notes follow the Lyceum API documentation, checked 6 October 2026.
Similar models
Other models in our catalogue for the same kind of work.