What is new in Qwen3.8 2.4T A95B
Modern agentic systems require extreme parameter scale to navigate multi-step code execution, autonomous tool orchestration, and deep reasoning across thousands of tokens. However, executing multi-trillion parameter dense models in production quickly becomes economically and computationally impractical. Qwen3.8 2.4T A95B addresses this bottleneck by introducing a massive 2.4 trillion total parameter Mixture-of-Experts (MoE) architecture designed specifically as an open-weight foundation for complex agentic workflows.
By activating only 95 billion parameters per token during inference, the model provides frontier-level reasoning depth while keeping per-step computational overhead aligned with sub-100B parameter footprints. This architecture allows AI-native product teams to execute complex autonomous loops without bearing the extreme latency and memory penalties associated with traditional dense architectures of comparable parameter counts.
- 2.4 trillion total parameters with 95 billion active parameters per token pass.
- Native 256K-token context window powered by a hybrid full and linear attention mechanism.
- Configurable inference-time reasoning controls for dynamic trade-offs between depth and latency.
- Immediate production availability on Serverless Inference without infrastructure provisioning.
The model is now deployed on Serverless Inference, giving development teams immediate access through an OpenAI-compatible API without requiring dedicated GPU cluster provisioning, custom kernel patching, or minimum spend commitments.
Mixture-of-experts architecture and parameters
The core engineering achievement of Qwen3.8 2.4T A95B lies in its fine-grained Mixture-of-Experts topology. Rather than relying on a small collection of coarse, high-capacity expert networks, the model distributes its 2.4 trillion parameters across 512 routed experts alongside a dedicated shared expert module. For every forward token pass, the routing mechanism dynamically selects 10 routed experts in parallel with the shared expert, engaging approximately 95 billion active parameters.
This fine-grained expert division prevents specialized domain knowledge from diluting across monolithic weights. The routing network learns discrete pathways for distinct computational requirements, such as deterministic syntax validation, algorithmic synthesis, and multi-lingual translation, activating only the precise sub-networks required for a given token sequence.
Hybrid backbone layer allocation
Underneath the MoE routing layer, the network operates on a 92-layer hybrid backbone designed to balance global contextual awareness with linear memory scaling. Standard transformer architectures incur quadratic memory expansion as context depth grows, creating acute Key-Value (KV) cache bottlenecks on high-concurrency systems. Qwen3.8 2.4T A95B decouples this relationship by interleaving full-attention layers with linear-attention recurrent states.
| Structural Dimension | Engineering Specification | Operational Impact |
|---|---|---|
| Total Parameters | 2.4 Trillion (2.4T) | Broad multi-domain pretraining capacity and representation capacity |
| Active Parameters | 95 Billion (95B) | Inference compute footprint comparable to medium dense models |
| Total Layers | 92 Backbone Layers | Deep feature transformation and multi-stage semantic extraction |
| Full Attention Layers | 23 Layers | Global cross-token attention for precise needle-in-haystack retrieval |
| Linear Attention Layers | 69 Gated DeltaNet Layers | Bounded recurrent state eliminating quadratic KV memory expansion |
| Expert Allocation | 512 Routed + 1 Shared | Fine-grained specialization with 10 routed experts activated per token |
With 23 full-attention layers interspersed across 69 Gated DeltaNet linear-attention layers, the model maintains full quadratic precision where global token relationships are essential, while delegating sequential state tracking to constant-memory recurrent layers. This architectural co-design prevents memory thrashing during intensive agentic coding tasks that generate extensive token sequences.
The 256K context window for agentic workflows
Agentic architectures fundamentally break the assumptions of traditional single-turn prompt-response systems. In production agents, the model must simultaneously hold complex system prompts, intermediate tool outputs, structured JSON schemas, API call histories, raw codebase directories, and multi-step reasoning traces. In conventional transformer designs, sustaining this state across hundreds of turns triggers severe memory pressure and latency degradation.
Qwen3.8 2.4T A95B provides a native 256K-token context window specifically engineered for long-horizon autonomous workflows. The 69 linear-attention layers replace standard unbound KV caching with a bounded recurrent state. As a result, memory consumption remains strictly controlled even as the prompt approaches the 256K boundary, allowing applications to ingest complete repositories or multi-hundred-page technical documentation without encountering Out-of-Memory (OOM) failures.
- Repository-scale context ingestion: Pass complete multi-file codebases and execution logs in a single context window.
- Bounded KV cache footprint: Linear-attention recurrence eliminates quadratic memory spikes during multi-turn agent execution.
- Zero prompt truncation: Maintain persistent system instructions and tool documentation across extended reasoning chains.
This design prevents the performance degradation commonly observed in dense models when processing massive documents. By maintaining a bounded recurrent state across the majority of its backbone, Qwen3.8 2.4T A95B delivers sustained token throughput regardless of context fill percentage.
Vendor-reported reasoning and coding capabilities
In benchmarks reported by the vendor (source: Qwen and NVIDIA), Qwen3.8 2.4T A95B exhibits top-tier performance across complex coding, mathematical reasoning, and multi-step tool orchestration evaluations. The model introduces configurable reasoning controls that allow developers to dictate the inference depth per request, modulating between fast deterministic code generation and extensive multi-step problem solving depending on workload requirements.
According to vendor data, the model demonstrates high instruction-following fidelity in autonomous coding benchmarks like CodeArena and complex multi-hop tool execution environments. On accelerated datacenter infrastructure, including NVIDIA GB300 NVL72 systems, the vendor reports peak hardware throughput exceeding 4,000 tokens per second per GPU in FP8 precision, with individual user generation speeds exceeding 350 tokens per second.
| Capability Metric | Vendor-Reported Value | Evaluation Context |
|---|---|---|
| Peak System Throughput | Over 4,000 tokens/sec/GPU | Vendor-reported benchmark on NVIDIA GB300 NVL72 in FP8 |
| Per-User Stream Speed | Over 350 tokens/sec | Vendor-reported single-user generation speed on Blackwell hardware |
| Reasoning Configuration | Low / High / Extra-High | Configurable inference-time reasoning effort per request |
| Context Retention | 256K native | Context window served on Serverless Inference for long-document extraction |
All benchmark figures and throughput metrics reflect vendor-reported evaluations and are provided as technical reference points for evaluating compute requirements. Production throughput on shared endpoints varies dynamically based on prompt length, batch concurrency, and active system load.
Pricing on Serverless Inference
The model is available through a direct pay-per-token plan on Serverless Inference, which avoids the capital inefficiency of maintaining dedicated GPU instances for intermittent agentic workloads. Billed units are tracked per million tokens: $2.50 per 1M input tokens, $0.63 per 1M cached input tokens, and $6.00 per 1M output tokens, as published in the catalogue record.
| Token Category | Unit Price (USD) | Billing Unit |
|---|---|---|
| Uncached Input Tokens | $2.50 | Per 1M tokens |
| Cached Input Tokens | $0.63 | Per 1M tokens (prompt cache hit) |
| Output Tokens | $6.00 | Per 1M generated tokens |
Prompt caching is automatically evaluated on recurring prompt prefixes, so repeated system prompts, large documentation contexts, and standard tool definitions are billed at the lower cached input rate shown above. This makes high-frequency agent polling significantly more cost-effective compared to traditional flat-rate token pricing models, shifting the overall cost per million tokens downward for production pipelines.
Data residency and hosting location: Qwen3.8 2.4T A95B is hosted in the EU, providing strict geographical data residency guarantees for teams routing sensitive production traffic.
How to call the Qwen3.8 2.4T A95B API
Integrating Qwen3.8 2.4T A95B into existing software pipelines requires only updating the base URL and specifying the model identifier. The endpoint provides full compatibility with standard OpenAI client libraries, allowing developers to reuse existing codebases without modifying serialization logic, tool schemas, or streaming parsers.
Python SDK integration
To call the model using the official Python client, configure the OpenAI instance to point to the Serverless Inference base URL and pass your API key as a standard Bearer token.
from openai import OpenAI
client = OpenAI(
base_url="https://api.lyceum.technology/openai/v1", api_key="lk_your_api_key_here"
)
response = client.chat.completions.create(
model="qwen/qwen3.8-2.4t-a95b",
messages=[
{
"role": "system",
"content": "You are an expert GPU runtime engineer.",
},
{
"role": "user",
"content": "Compare Gated DeltaNet and multi-head attention memory use.",
},
],
temperature=0.2,
max_tokens=1024,
)
print(response.choices[0].message.content)Direct HTTP curl execution
For direct HTTP integration in containerized microservices or command-line scripts, submit a POST request to the chat completions endpoint using standard JSON payloads:
curl https://api.lyceum.technology/openai/v1/chat/completions \
-H "Authorization: Bearer lk_your_api_key_here" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen/qwen3.8-2.4t-a95b",
"messages": [
{"role": "user", "content": "Write a CUDA kernel launch config check in C++."
}],
"temperature": 0.1
}'The API supports token streaming for real-time user interfaces, structured JSON output, and standard function calling definitions out of the box.
Where the model fits in the catalogue
Within the Model Library, Qwen3.8 2.4T A95B sits as the high-capacity anchor for multi-step reasoning and deep contextual synthesis. For workloads that do not require 2.4 trillion parameters of general knowledge, smaller variants such as Qwen3.8 27B or Qwen3.8-Flash-Next offer lower latency profiles for deterministic processing and routing; check each model's own catalogue record for its current per-token rates.
| Model Offering | Parameter Scale | Context Window | Input / Output (per 1M) | Primary Workload Profile |
|---|---|---|---|---|
| Qwen3.8 2.4T A95B | 2.4T MoE (95B active) | 256K | $2.50 / $6.00 | Complex agentic loops, deep code synthesis, long document analysis |
| Qwen3.8 27B | 27B Dense | 256K | See catalogue record | High-throughput structured extraction, fast code refactoring |
| Qwen3.8-Flash-Next | Flash MoE | 256K | See catalogue record | Real-time agentic routing, lightweight tool execution |
Operational terms on Serverless Inference are completely self-serve. There are no minimum monthly spend commitments, base subscription fees, or capacity reservation contracts required. Requests are metered purely on processed token volume with zero idle compute waste.
Service availability specifications: Serverless Inference and Smart Routing operate as self-serve per-token endpoints and carry no formal Service Level Agreement (SLA), availability tier, uptime target, or service credit mechanism. Teams that require contractually guaranteed GPU capacity, dedicated instance isolation, or formal enterprise SLAs should deploy via Dedicated Inference or On-demand GPU VM infrastructure instead.
For AI-native teams building the next generation of autonomous tools, deploying Qwen3.8 2.4T A95B via Serverless Inference delivers immediate access to frontier-scale open weights through a clean, OpenAI-compatible API without the burden of infrastructure management.
Running Qwen3.8 2.4T A95B in Claude Code
Lyceum exposes an Anthropic-compatible endpoint alongside the OpenAI-compatible one, so Claude Code can be pointed at Qwen3.8 2.4T A95B without a separate adapter. The Lyceum CLI does the setup for you: lyceum code launches Claude Code preconfigured in its own profile, which leaves an existing Anthropic login untouched. To wire it up by hand, set three environment variables and start Claude Code:
export ANTHROPIC_BASE_URL=https://api.lyceum.technology/anthropic
export ANTHROPIC_AUTH_TOKEN=lk_your_api_key_here
export ANTHROPIC_MODEL=qwen/qwen3.8-2.4t-a95b
claudeUse ANTHROPIC_AUTH_TOKEN rather than ANTHROPIC_API_KEY. Both authenticate, but ANTHROPIC_API_KEY makes Claude Code ask for approval of a custom key on first run, which fails outright in non-interactive use such as claude -p or CI with "Not logged in - Please run /login". The tier variables ANTHROPIC_DEFAULT_OPUS_MODEL, ANTHROPIC_DEFAULT_SONNET_MODEL and ANTHROPIC_DEFAULT_HAIKU_MODEL map the Claude tier names in the /model menu onto Lyceum models, so picking a tier selects the model you mapped to it.