Moonshot
Kimi K2.7 Code
A coding model built on Kimi K2.6 for long software engineering tasks, which Moonshot says uses about 30% fewer thinking tokens than K2.6.
- Context window
- 256K tokens
- Accepts
- Text and images
- Input per 1M tokens
- $1.25
- Output per 1M tokens
- $4.50
Send your first request
Set LYCEUM_API_KEY to your API key, then run this request. Usage is billed to your account.
- Base URL
https://api.lyceum.technology/openai/v1- Model ID
moonshotai/kimi-k2.7-code- Request fields
model, messages, stream, max_tokens, tools, tool_choice
Documented request fields for this endpoint. See the documentation for model-specific controls and limits.
app.py · Install openai and set LYCEUM_API_KEY
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["LYCEUM_API_KEY"],
base_url="https://api.lyceum.technology/openai/v1",
)
response = client.chat.completions.create(
model="moonshotai/kimi-k2.7-code",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)app.ts · Install openai and set LYCEUM_API_KEY in Node.js
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.LYCEUM_API_KEY,
baseURL: "https://api.lyceum.technology/openai/v1",
});
const response = await client.chat.completions.create({
model: "moonshotai/kimi-k2.7-code",
messages: [{ role: "user", content: "Hello!" }],
});
console.log(response.choices[0]?.message.content);Terminal · Set LYCEUM_API_KEY in your terminal
curl https://api.lyceum.technology/openai/v1/chat/completions \
-H "Authorization: Bearer $LYCEUM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "moonshotai/kimi-k2.7-code",
"messages": [
{"role": "user", "content": "Hello!"}
]
}'What it costs and how fast it runs
You pay only for the tokens you use. Speed figures come from our own status checks.
Token pricing
| USD per million tokens | |
|---|---|
| Input | $1.25 |
| Cached input | $0.31 |
| Output | $4.50 |
No base fee. Caching applies only where listed. Check the model's processing region before sending data.
Speed on Lyceum
| Last 7 days | |
|---|---|
| Time to first token | 324 ms |
| Output speed | 109 tokens/s |
Median over the last 7 days from our status checks: short requests. Measured 6 October 2026, 23:08 UTC. Reasoning tokens count towards both figures. Live status and availability
The model in detail
Checked against the sources below. Anything they don't state is left out.
Capabilities
- Input and output
- Text and images in, text out
- Context window
- 256K tokens
- Reasoning
- On by default
- Tool calling
- Supported
Architecture and licence
- Total parameters
- 1T
- Active parameters
- 32B
- Architecture
- Mixture of experts
- Licence
- Modified MIT License
- Released
- 12 June 2026
Availability
- API model ID
moonshotai/kimi-k2.7-code- Provider
- Moonshot
- Processing region
- EU-hosted
- Input and output types: official model card, checked 10 September 2026.
- Capabilities and request fields: Lyceum API documentation, checked 9 September 2026.
- Parameters, licence and release date: Moonshot’s own sources, including its model card and announcement, checked 29 September 2026.
Common questions about Kimi K2.7 Code
Short answers about the model ID, price, limits and behaviour.
What is the model ID for Kimi K2.7 Code?
Use moonshotai/kimi-k2.7-code as the model value. The OpenAI-compatible base URL is https://api.lyceum.technology/openai/v1. You authenticate with your Lyceum API key as a Bearer token.
How much does Kimi K2.7 Code cost?
Kimi K2.7 Code costs $1.25 per 1M input tokens, $0.31 per 1M cached input tokens and $4.50 per 1M output tokens. Prices are in US dollars. Billed per token. No base fee.
What is the context window of Kimi K2.7 Code?
Kimi K2.7 Code has a 256K tokens context window. One request can generate up to 65,536 tokens and run for up to 300 seconds. Set stream: true for long outputs.
Can Kimi K2.7 Code read images?
Yes. Send images as OpenAI-style image_url content parts, from a public URL or a base64 data URL. It also reads PDFs, sent as a data:application/pdf;base64 URL.
Can I turn off reasoning for Kimi K2.7 Code?
No. Reasoning is always on for Kimi K2.7 Code. The request is accepted, but the setting is ignored. The reasoning trace counts as output tokens, so set max_tokens high enough for both the reasoning and the answer.
Does Kimi K2.7 Code support tool calling?
Yes. Send tools in the standard OpenAI format. Automatic tool calls, parallel calls and tool-result round trips work.
Does Kimi K2.7 Code support prompt caching?
Yes, and it is automatic. When requests share an identical start, the cached part is billed at $0.31 per 1M tokens instead of $1.25. Keep the stable part of your prompt at the front. A cache hit is best effort, so read usage.prompt_tokens_details.cached_tokens, which is missing on a miss.
Where does Kimi K2.7 Code run, and is my data kept?
Kimi K2.7 Code runs on EU-hosted infrastructure. Prompts and outputs are processed, not stored, and never used for training.
How do I call Kimi K2.7 Code with the OpenAI SDK?
Point the OpenAI SDK at https://api.lyceum.technology/openai/v1, use your Lyceum API key and set model="moonshotai/kimi-k2.7-code". Chat completions, streaming, tool calling and usage accounting work as they do with OpenAI.
Behaviour notes follow the Lyceum API documentation, checked 6 October 2026.
Similar models
Other models in our catalogue for the same kind of work.