Z.ai
GLM-5.2
Z.ai’s model for long agentic coding tasks, with a 1M-token context and a choice of how long it thinks.
- Context window
- 1M tokens
- Accepts
- Text
- Input per 1M tokens
- $1.50
- Output per 1M tokens
- $4.50
Send your first request
Set LYCEUM_API_KEY to your API key, then run this request. Usage is billed to your account.
- Base URL
https://api.lyceum.technology/openai/v1- Model ID
z-ai/glm-5.2- Request fields
model, messages, stream, max_tokens, tools, tool_choice
Documented request fields for this endpoint. See the documentation for model-specific controls and limits.
app.py · Install openai and set LYCEUM_API_KEY
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["LYCEUM_API_KEY"],
base_url="https://api.lyceum.technology/openai/v1",
)
response = client.chat.completions.create(
model="z-ai/glm-5.2",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)app.ts · Install openai and set LYCEUM_API_KEY in Node.js
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.LYCEUM_API_KEY,
baseURL: "https://api.lyceum.technology/openai/v1",
});
const response = await client.chat.completions.create({
model: "z-ai/glm-5.2",
messages: [{ role: "user", content: "Hello!" }],
});
console.log(response.choices[0]?.message.content);Terminal · Set LYCEUM_API_KEY in your terminal
curl https://api.lyceum.technology/openai/v1/chat/completions \
-H "Authorization: Bearer $LYCEUM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "z-ai/glm-5.2",
"messages": [
{"role": "user", "content": "Hello!"}
]
}'What it costs and how fast it runs
You pay only for the tokens you use. Speed figures come from our own status checks.
Token pricing
| USD per million tokens | |
|---|---|
| Input | $1.50 |
| Cached input | $0.38 |
| Output | $4.50 |
No base fee. Caching applies only where listed. Check the model's processing region before sending data.
Speed on Lyceum
| Last 7 days | |
|---|---|
| Time to first token | 385 ms |
| Output speed | 51 tokens/s |
Median over the last 7 days from our status checks: short requests. Measured 6 October 2026, 22:38 UTC. Reasoning tokens count towards both figures. Live status and availability
The model in detail
Checked against the sources below. Anything they don't state is left out.
Capabilities
- Input and output
- Text in, text out
- Context window
- 1M tokens
- Reasoning
- On by default
- Tool calling
- Supported
Architecture and licence
- Total parameters
- 744B
- Active parameters
- 40B
- Architecture
- Mixture of experts
- Licence
- MIT
- Released
- 16 June 2026
Availability
- API model ID
z-ai/glm-5.2- Provider
- Z.ai
- Processing region
- EU-hosted
- Input and output types: official model card, checked 10 September 2026.
- Capabilities and request fields: Lyceum API documentation, checked 9 September 2026.
- Parameters, licence and release date: Z.ai’s own sources, including its model card and announcement, checked 29 September 2026.
Common questions about GLM-5.2
Short answers about the model ID, price, limits and behaviour.
What is the model ID for GLM-5.2?
Use z-ai/glm-5.2 as the model value. The OpenAI-compatible base URL is https://api.lyceum.technology/openai/v1. You authenticate with your Lyceum API key as a Bearer token.
How much does GLM-5.2 cost?
GLM-5.2 costs $1.50 per 1M input tokens, $0.38 per 1M cached input tokens and $4.50 per 1M output tokens. Prices are in US dollars. Billed per token. No base fee.
What is the context window of GLM-5.2?
GLM-5.2 has a 1M tokens context window. One request can generate up to 65,536 tokens and run for up to 300 seconds. Set stream: true for long outputs.
Can GLM-5.2 read images?
No, not today. The request succeeds, but the image is ignored. Pick a model that reads images if you need that.
Can I turn off reasoning for GLM-5.2?
Yes. Reasoning is on by default. Send reasoning_effort: "none" in the request, or call z-ai/glm-5.2-instant, the same model at the same price without a reasoning phase. The reasoning trace counts as output tokens when it is on, so set max_tokens high enough for both the reasoning and the answer.
Does GLM-5.2 support tool calling?
Yes. Send tools in the standard OpenAI format. Automatic tool calls, parallel calls and tool-result round trips work. It does not enforce tool_choice: "required" or a named function. It can answer in plain text with no tool calls, so check message.tool_calls and fall back when it is empty.
Does GLM-5.2 support prompt caching?
Yes, and it is automatic. When requests share an identical start, the cached part is billed at $0.38 per 1M tokens instead of $1.50. Keep the stable part of your prompt at the front. A cache hit is best effort, so read usage.prompt_tokens_details.cached_tokens, which is missing on a miss.
Where does GLM-5.2 run, and is my data kept?
GLM-5.2 runs on EU-hosted infrastructure. Prompts and outputs are processed, not stored, and never used for training.
How do I call GLM-5.2 with the OpenAI SDK?
Point the OpenAI SDK at https://api.lyceum.technology/openai/v1, use your Lyceum API key and set model="z-ai/glm-5.2". Chat completions, streaming, tool calling and usage accounting work as they do with OpenAI.
Behaviour notes follow the Lyceum API documentation, checked 6 October 2026.
Similar models
Other models in our catalogue for the same kind of work.