What is new in Qwen3.8 27B

Qwen3.8 27B is now available on Serverless Inference. Released by the Qwen team on August 14, 2026, this release packages a 27-billion-parameter dense multimodal architecture under the Apache 2.0 license. For AI-native product teams building autonomous agent loops, code generation pipelines, and complex document workflows, this model provides frontier-grade reasoning and vision understanding in a compact parameter footprint.

While recent flagship models have scaled parameter counts into hundreds of billions or trillions using sparse Mixture-of-Experts (MoE) designs, dense models in the 20B to 30B class remain essential for predictable serving economics. They deliver high per-stream throughput, eliminate cross-node routing overhead, and fit into standard hardware configurations while retaining the architectural depth necessary for complex tool use and multi-step task execution.

  • Architecture: 27B dense multimodal causal language model with integrated vision encoder
  • Context window: 256K native context length (262,144 tokens) for long-document parsing and extended agent context
  • Licensing: Apache 2.0 open-weight distribution allowing unconstrained commercial deployment and fine-tuning
  • Acceleration: Native Multi-Token Prediction (MTP) draft head support inside the checkpoint for speculative decoding

Engineers can query Qwen3.8 27B immediately using standard OpenAI SDK clients without managing GPU nodes, scheduling container clusters, or handling CUDA memory fragmentation. The model sits alongside the wider selection in the Text Models Library, providing instant API access for production workloads.

Architecture and 256K context window

The core engineering highlight of Qwen3.8 27B is its hybrid attention mechanism. Rather than applying standard quadratic multi-head attention across all transformer blocks, the model uses a 64-layer architecture where only 16 layers execute full gated attention (full_attention_interval: 4). The remaining 48 layers run linear attention via Gated DeltaNet with a constant recurrent state.

Architectural ParameterSpecification ValueEngineering Impact
Total Parameters27 Billion (Dense)Compact footprint with no expert routing overhead
Transformer Layers64 total layers (16 Full Attention, 48 Linear)Hybrid design cutting linear KV cache growth
Hidden Dimension5120, with a 17,408 feed-forward intermediate dimensionMaintains representational capacity across reasoning tasks
Native Context Length262,144 tokens (256K), extensible to 1MSupports large codebase ingestion and multi-turn traces
Speculative DecodingIntegrated MTP draft headServing engines expose it via EAGLE-style speculative decoding

This 1:3 ratio between full attention and linear recurrence significantly reduces memory pressure during long-context processing. In standard architectures, key-value (KV) cache memory scales linearly with sequence length across every layer, frequently triggering Out-of-Memory (OOM) failures at 128K or 256K tokens unless heavy quantization or context offloading is applied. By bounding KV cache growth to 16 attention layers, Qwen3.8 27B maintains stable memory overhead even when processing massive document batches.

Regarding infrastructure compliance and data location: hosting region and EU residency are not supplied for this model in the catalogue. Data residency is treated as a per-model specification rather than an unverified platform-wide generalization. Developers requiring verified EU data residency should consult the dedicated model entry in the Model Library before routing production workloads.

Vendor-reported capabilities and benchmarks

To evaluate Qwen3.8 27B against existing open-weight and proprietary baselines, we look at the standardized benchmark results published at release. Every figure below is a vendor-reported score printed on the Qwen model card and specific to the harness Qwen used: the card's own footnotes describe QwenSWEBench as an in-house coding benchmark and CoWorkBench as an in-house cowork benchmark, and several coding runs use the Claude Code harness. None of them are measured on Serverless Inference.

BenchmarkQwen3.8-27B ScoreEvaluation DomainBenchmark Significance
QwenSWEBench79.0%Software EngineeringEnd-to-end bug fixing and pull request resolution
LiveCodeBench v690.3%Competitive CodingContamination-free algorithmic programming
SWE-bench Pro61.7%Agentic CodingResolving complex open-source issue threads
GPQA Diamond89.2%Graduate ScienceExpert-level physics, chemistry, and biology reasoning
BullshitBench v278.0%Nonsense DetectionIdentifying invalid premises and resisting hallucinations
JobBench33.4%Professional Tool UseMulti-step office workflows and enterprise tool execution
CoWorkBench70.7%Long-Horizon TasksProductivity tasks across law, finance, and engineering

The benchmark profile demonstrates substantial improvements in agentic autonomy and multi-step tool orchestration. For engineering teams evaluating models for agentic coding and automated CI/CD triage, the 61.7% score on SWE-bench Pro and 90.3% on LiveCodeBench v6 place this 27B model ahead of several previous-generation dense architectures, though the SWE-bench Pro run used the Claude Code harness at temperature 1.0 with a 256K context window, so cross-model comparisons shift with the setup.

Managing built-in reasoning traces

Qwen3.8 27B incorporates built-in reasoning capabilities. Qwen describes the model as having flexible thinking control: thinking mode is on by default and can be disabled per request, and its chat template opens every assistant turn with a <think> block, so reasoning tokens precede the final answer unless a reasoning parser strips them. While this internal reasoning improves accuracy on mathematical proofs, algorithmic code synthesis, and multi-file debugging, production systems often require fine-grained control over when reasoning tokens are generated.

Developers can manage reasoning behavior dynamically on a per-request basis using the chat_template_kwargs parameter in the API payload. Passing {"enable_thinking": false} turns off the internal reasoning pass, causing the model to generate direct completions without <think> blocks. This eliminates extra latency and output token charges for simple classifications, formatting tasks, or high-volume summarization.

  • Default Thinking Mode: {"enable_thinking": true} (default). Emits full reasoning traces before the response. Recommended sampling from generation_config.json: temperature=1.0, top_p=0.95, top_k=20.
  • Direct Instruct Mode: {"enable_thinking": false} makes the model answer directly without thinking. Lower the temperature for deterministic formatting and classification work.
  • Reasoning Effort: reasoning depth is tunable with the reasoning_effort parameter, which accepts xhigh (default), medium, and low, trading chain-of-thought depth against latency and cost.
  • Preserve Thinking: Qwen states that reasoning context from historical messages is retained via preserve_thinking, which matters for multi-turn agent histories.

Note that on Serverless Inference, all requests are bounded by a maximum generation timeout of 300 seconds. For complex multi-step reasoning chains spanning tens of thousands of tokens, ensure your client-side HTTP timeouts accommodate extended generation times or stream responses incrementally.

Per-token pricing and prompt caching

Serverless Inference bills Qwen3.8 27B strictly on actual token consumption without base subscription fees, minimum hourly GPU commitments, or idle infrastructure charges. The standard list price is $0.40 per 1 million input tokens and $2.40 per 1 million output tokens.

For applications with large, repetitive prompt structures, such as agent system prompts, OpenAPI schemas, coding repository indexes, or few-shot exemplars, Serverless Inference provides automatic prefix prompt caching. When an incoming request shares a common token prefix with recent requests, the cached input segment is billed at $0.10 per 1 million tokens instead of the standard input rate.

Token CategoryStandard Price (per 1M Tokens)Effective Rate with Prompt CachingBilling Mechanism
Input Tokens (Uncached)$0.40$0.40Metered per prompt token processed
Cached Input Tokens$0.10$0.10Reduced rate on repeated prompt prefixes
Output Tokens$2.40$2.40Metered per completion token generated

In agentic architectures where an orchestration loop repeatedly sends the same large system prompt and tool definitions across many turns, prompt caching shifts the unit economics. Monitoring your cost per million tokens across cached and uncached streams enables accurate FinOps modeling as request volume scales.

Where it fits in the model catalogue

Selecting the right model requires balancing parameter capacity, latency constraints, and operational cost. Within the Serverless Inference catalogue, Qwen3.8 27B occupies the high-efficiency reasoning tier, positioned between lightweight 9B to 32B instruction models and massive trillion-parameter MoE architectures.

ModelArchitectureContext WindowInput / 1M TokensOutput / 1M Tokens
Qwen3.8 27B27B Dense (Multimodal)256K$0.40 ($0.10 cached)$2.40
Qwen3-32B32B Dense (Text/Code)128K$0.10$0.30
Qwen3-30B-A3B30B MoE (3B active)256K$0.10$0.30
GLM-5.3MoE Flagship1M$1.75 ($0.44 cached)$4.50

Compared to the sparse Qwen3-30B-A3B, which activates only 3 billion parameters per token and optimizes for raw throughput on simple tasks, Qwen3.8 27B engages its full 27-billion dense parameter set on every forward pass. This gives it significantly stronger problem-solving capabilities on multi-step reasoning and deep code analysis. Conversely, when compared to flagship MoEs like GLM-5.3 ($1.75 per 1M input tokens), Qwen3.8 27B delivers competitive agentic performance at roughly one-fourth the input price.

Deploying on Serverless Inference

Deploying Qwen3.8 27B in production requires no bespoke orchestration frameworks. Serverless Inference exposes a standard OpenAI-compatible API surface, allowing engineering teams to route traffic by updating their client base URL and model string while preserving existing streaming, tool-calling, and error-handling logic.

  1. Configure your API key: Generate an API key (lk_...) from the dashboard.
  2. Set the base URL: Point your client at the OpenAI-compatible base URL published in the Serverless Inference docs.
  3. Specify the model identifier: Pass qwen/qwen3.8-27b in your request body.
  4. Configure template options: Pass enable_thinking and reasoning_effort inside extra_body if you wish to adjust default reasoning behavior.

The following Python example shows how to execute a completion request using the official OpenAI SDK:

  • Client Initialization: Set base_url to the OpenAI-compatible Serverless Inference endpoint and api_key="lk_...".
  • Chat Completion Call: Pass model="qwen/qwen3.8-27b" and your prompt messages array.
  • Reasoning Control: Include extra_body={"chat_template_kwargs": {"enable_thinking": True}} to control reasoning traces.
  • Streaming & Metrics: Set stream=True and capture usage tokens directly from stream completion chunks.

Operational note on service level agreements: Serverless Inference is a self-serve, pay-per-token product and carries no SLA, availability tier, uptime target, or service credit commitments. Production engineering teams with strict contractual availability requirements should speak with our infrastructure team regarding dedicated GPU capacity.

  • Review the exact model specifications in the Serverless Inference documentation.
  • Benchmark your prompt prefix cache hit rates using the $0.10/1M cached input rate.
  • Deploy workloads directly to Serverless Inference to scale without infrastructure management.

Running Qwen3.8 27B in Claude Code

Lyceum exposes an Anthropic-compatible endpoint alongside the OpenAI-compatible one, so Claude Code can be pointed at Qwen3.8 27B without a separate adapter. The Lyceum CLI does the setup for you: lyceum code launches Claude Code preconfigured in its own profile, which leaves an existing Anthropic login untouched. To wire it up by hand, set three environment variables and start Claude Code:

export ANTHROPIC_BASE_URL=https://api.lyceum.technology/anthropic
export ANTHROPIC_AUTH_TOKEN=lk_your_api_key_here
export ANTHROPIC_MODEL=qwen/qwen3.8-27b
claude

Use ANTHROPIC_AUTH_TOKEN rather than ANTHROPIC_API_KEY. Both authenticate, but ANTHROPIC_API_KEY makes Claude Code ask for approval of a custom key on first run, which fails outright in non-interactive use such as claude -p or CI with "Not logged in - Please run /login". The tier variables ANTHROPIC_DEFAULT_OPUS_MODEL, ANTHROPIC_DEFAULT_SONNET_MODEL and ANTHROPIC_DEFAULT_HAIKU_MODEL map the Claude tier names in the /model menu onto Lyceum models, so picking a tier selects the model you mapped to it.