AI This article was created with the help of AI.

An empty reply is not always an error

A coding assistant can show a blank panel even when the request returns HTTP 200 and reports generated tokens. Start with the raw response. The client may still be waiting for answer text, the output budget may have run out, or the response may need handling that the client has not completed.

HTTP success means the server returned a response, not that it produced a useful final answer. Check finish_reason alongside the message fields and usage. A token count proves that tokens were reported; it does not prove an answer reached the editor or establish the final invoice charge.

This guide covers 2 useful checks: output-budget exhaustion and client response handling. They are not an exhaustive list of causes. Tool calls, stop settings, refusals, interrupted streams and provider-side parsing issues can also leave no visible answer.

  • An empty content field can coexist with generated reasoning tokens
  • A length finish reason plus reasoning and no content is consistent with the output budget running out before the answer
  • If the same response contains answer text but the panel is blank, inspect client parsing, streaming and display handling
  • A tool_calls finish reason can require a tool result before a final answer appears

This guide explains how to investigate a blank reply on Lyceum Serverless Inference. Read the complete response, then test a larger output budget or an available instant variant where appropriate. The decision table below separates budget problems from client problems.

Where the reasoning text actually goes

Some serving stacks expose reasoning separately from final answer text. Current vLLM documentation names that extra field reasoning. Other providers or versions may use reasoning_content or expose no reasoning text. Read the response contract for the endpoint you are using.

  • Non-streaming returns one completed response. If content is empty, it will not fill in later inside that response object
  • Streaming can send reasoning before answer deltas. A client displaying only content may look idle during that phase
  • vLLM changed reasoning_content to reasoning. Clients must read the field their endpoint actually returns; field-name differences do not by themselves explain missing final content

For integration checks beyond reasoning, see what breaks when you switch models on an OpenAI-compatible API. Compare the request sent by the editor with your direct test, including output limits, tools and stop settings. A successful test with different settings does not isolate the editor as the cause.

When the output budget runs out first

For the GLM-5.2 route tested here, max_tokens limits the completion budget, including reasoning. If generation reaches that limit before producing answer text, finish_reason can be length with content null. The response can still report completion tokens.

In a live check on 17 September 2026, z-ai/glm-5.2 received the decimal-comparison prompt below with max_tokens set to 20. It returned finish_reason length, content null and a populated reasoning field. Usage reported 20 completion tokens, all classified as reasoning tokens. This was a response-shape check, not a benchmark or invoice audit.

Allow room for reasoning and the final answer. The amount needed depends on the model and task; a single fixed cap is not a guarantee.

A client configured with a small output limit can expose this behaviour when you change models. Both open-weight and closed models may reason before answering. Check the actual request limit rather than assuming it from the tool or model label. Output limits and context-window limits are separate constraints.

A second request with the same prompt and a 2,000-token cap returned answer text and finish_reason stop. It reported 551 completion tokens, including 381 reasoning tokens. The result supports the budget diagnosis for these requests; it does not guarantee that 2,000 tokens will cover every task.

When the client cannot render the field

If a captured response contains final answer text but the editor shows none, investigate the client or intermediary. Hiding a reasoning field alone does not explain why populated content is missing. For streaming, check whether the client waited for the answer and handled all events. If the server itself returns empty content with finish_reason stop, inspect stop settings, parser behaviour and other message fields.

  • VS Code supports custom model providers. Its model configuration includes reasoning capabilities; a missing UI control alone does not establish why a reply is blank
  • The Thinking Effort control depends on model support and configuration. The VS Code reference includes supportsReasoningEffort for custom models
  • Cursor documents its OpenAI API-key option for standard, non-reasoning chat models. Verify your custom endpoint and model in the installed client; do not assume every reasoning response format is supported

For configuration, see the opencode custom provider setup or the current Claude Code setup. Lyceum recommends its lyceum code launcher for Claude Code. These are separate client routes to test, not proof that a failing Cursor or Copilot request has been fixed.

Choosing an instant variant instead

Lyceum documents its GLM-5.2 instant variant as the same model with thinking disabled. Thinking controls vary across model families and serving configurations. vLLM documents such controls, but its Qwen configuration is not proof of how every hosted variant is implemented.

The live /openai/v1/models response listed these 5 instant IDs on 17 September 2026. Listing confirms roster presence, not a successful inference test for every variant. The GLM-5.2 instant example below was tested directly.

API model stringReview evidence
z-ai/glm-5.2-instantListed; visible content returned in the test below
z-ai/glm-5.3-instantListed; not individually inference-tested in this review
z-ai/glm-5.3-flash-instantListed; not individually inference-tested in this review
qwen/qwen3.8-27b-instantListed; not individually inference-tested in this review
qwen/qwen3.8-flash-next-instantListed; not individually inference-tested in this review
SymptomFixWhen to choose it
length, no content, populated reasoningIncrease the output budget and rerunThe response is consistent with the cap interrupting reasoning
Answer text in the captured response, blank editorInspect client parsing and stream handlingThe answer exists but is not shown
stop, empty contentInspect stop settings, tool_calls and provider parsingThis does not prove a display problem
A low client cap cannot be changedTest an available instant variantAccept any quality trade-off and verify the result

Increase the budget when reasoning helps the task. Consider an instant variant when a client imposes a small limit or the task does not need a thinking phase. Neither change repairs every integration issue. Check current model rates and measure total tokens before claiming a saving.

Diagnosing it with one request

One direct, non-streaming request is a useful starting point. It shows what the endpoint returns without the editor. Follow-up requests may be needed to isolate the cause.

The example deliberately sets max_tokens to 20. This reproduced the budget symptom in the check above, but output can vary. Set LYCEUM_API_KEY in your shell, then send the request and inspect the complete response.

curl https://api.lyceum.technology/openai/v1/chat/completions \
  -H "Authorization: Bearer $LYCEUM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "z-ai/glm-5.2",
    "messages": [{"role": "user", "content": "Which is greater, 9.11 or 9.8?"}],
    "max_tokens": 20
  }'

Read the complete message, finish_reason and usage. Reasoning text shows that reasoning was returned; it does not establish that the final answer was generated or that the client received it.

Use these checks to choose the next step:

  1. length with empty content and populated reasoning: increase the cap and compare the result. Check the endpoint’s output-limit parameter if it differs from max_tokens.
  2. Populated content in the captured response but a blank panel: inspect the editor or intermediary. Compare equivalent requests before attributing the cause.
  3. Empty content with stop, or both text fields empty: inspect tool_calls, refusal and stop settings. Preserve the response and request ID for support; remove credentials and sensitive prompt text before sharing.

The OpenAI Python client accepts a custom base URL and provider key. The example below prints both possible reasoning fields, the answer and tool calls. Those extra fields depend on the provider; do not treat an absent field as proof that no reasoning occurred.

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.lyceum.technology/openai/v1",
    api_key=os.environ["LYCEUM_API_KEY"],
)
response = client.chat.completions.create(
    model="z-ai/glm-5.2",
    messages=[{"role": "user", "content": "Which is greater, 9.11 or 9.8?"}],
    max_tokens=20,
)
choice = response.choices[0]
print("finish_reason:", choice.finish_reason)
print("reasoning:", getattr(choice.message, "reasoning", None))
print("reasoning_content:", getattr(choice.message, "reasoning_content", None))
print("content:", choice.message.content)
print("tool_calls:", choice.message.tool_calls)
print("usage:", response.usage)

For a team rollout, compare task quality, total billed tokens and response time on your own prompts. Disabling thinking can reduce reasoning output, but it does not guarantee lower total cost or equivalent answers.

If you are comparing open models for coding agents, run this diagnostic against each candidate before committing, because reasoning behaviour differs per model family. Our comparison of the best open model APIs for agentic coding covers the shortlist itself.

To compare variants on the same endpoint, change z-ai/glm-5.2 to z-ai/glm-5.2-instant. Keep the prompt and cap unchanged for this small diagnostic. This tests the response shape for that variant; it does not validate the editor integration.

Once you know which cause you hit, verify the fix rather than assuming it. Rerun the same call with the model string changed to the instant variant:

curl https://api.lyceum.technology/openai/v1/chat/completions \
  -H "Authorization: Bearer $LYCEUM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "z-ai/glm-5.2-instant",
    "messages": [{"role": "user", "content": "Which is greater, 9.11 or 9.8?"}],
    "max_tokens": 20
  }'
  • In the 17 September check, the instant variant returned visible answer text and no reasoning text at a 20-token cap
  • Usage reported 20 completion tokens and zero reasoning tokens for that request
  • finish_reason was still length: the visible answer was truncated. Increase the cap to obtain a complete response

Keep the diagnostic alongside the request settings and client version. Rerun it when you change models or client configuration, and compare the raw response with what the editor displays.

An instant variant changes model behaviour, not just what the interface displays. Test whether it still handles your task well. For work that benefits from reasoning, try a sufficient output budget before changing variants.

The 20-token and 2,000-token checks above demonstrate one budget-related failure and one completed answer. They are not timing measurements, quality benchmarks or universal output-budget recommendations.

Choose the fix from the evidence: a cap reached during reasoning calls for a budget check; an answer present in the captured response calls for a client check. An empty response with stop needs further investigation.

Get an API key in the Lyceum dashboard and try it on your own repository.