DeepSeek-V4.1-Flash is live on Lyceum Serverless Inference

DeepSeek-V4.1-Flash is live on Lyceum Serverless Inference: this page covers what changed, what it costs per token and the exact request to send. The model is pre-hosted on the shared endpoint, billed per token, with no minimum commitment, and you call it with the model string deepseek/deepseek-v4.1-flash. The price is $0.50 per 1M input tokens, $0.13 per 1M cached input tokens and $1.50 per 1M output tokens.

The sections below cover DeepSeek's reported architecture and benchmarks, Lyceum's current token prices, and working text-chat requests. Upstream specifications and vendor benchmark results do not establish every capability or limit of the hosted endpoint.

  • Model string: deepseek/deepseek-v4.1-flash
  • Price: $0.50 input / $0.13 cached input / $1.50 output per 1M tokens (USD, checked 17 September 2026)
  • Billing: per token, no minimum commitment
  • Hosted example: text input and text output. Image support is an upstream capability; verify it on your chosen Lyceum route.
  • Weights: MIT License, per DeepSeek's model card

Hosting region is a per-model and per-route fact. The live model list and price book used for this review do not establish residency. If you need EU processing, confirm the route and contractual data-handling terms with Lyceum before sending that workload.

Encoder-decoder MoE: 8B active for input, 16B for output

The architecture figures in this section come from DeepSeek's model card. DeepSeek-V4.1-Flash is a Mixture-of-Experts model with a 552B backbone. The unusual part is the activation split: 8B parameters active per token during prefill and 16B during decode. That asymmetry comes from the Causal Encoder-Decoder (CED) architecture, a 40-layer Transformer organised as a 20-layer causal encoder followed by a 20-layer decoder. The decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states, which DeepSeek designed to reduce prefill computation.

  • Backbone: 552B parameters, MoE
  • Active parameters: 8B per token in prefill, 16B in decode
  • Layout: 40 layers, 20-layer causal encoder plus 20-layer decoder
  • MoE routing: 1 shared and 384 routed experts per layer, 6 routed experts active per token
  • Conditional memory: Engram, 196B parameters, sparsely accessed via token-based lookup
  • Decoding: DSpark speculative decoding, semi-autoregressive draft generation with confidence-scheduled verification

DeepSeek designed this model for input-heavy agent workloads with long prompts, tool results and retrieved context. The upstream model processes images and text and generates text. Its model card reports a 45T-token training corpus, 64K training sequences and later extension to 1M context. MIT-licensed weights support inspection and self-hosting. They do not expose Lyceum's serving runtime or prove the exact precision and configuration used by a managed route.

A quarter of V4-Flash's KV cache per token

DeepSeek reports a global KV cache of 890 bytes per token, roughly 1/4 of DeepSeek-V4-Flash. Sliding-window attention (SWA) Bounded Replay reconstructs missing sliding-window KV states by replaying only the most recent n_win tokens, which avoids persisting SWA KV to SSD and cuts the persistent KV cache footprint to roughly 1/8 of V4-Flash's. DeepSeek's own figure-1 caption puts the reduction at approximately 4-fold against V4-Flash and 437-fold against DeepSeek-V1.

Compressed Sparse Attention 2 assigns each attention layer one of three static modes, Full, Reindex or Reuse, so main KV and indexer K are shared across layers and Top-K sparse-attention indices are reused. Combined with FP4 main KV caching in E2M1 format with one E4M3 scale per 16 channels, that is what brings the global cache down to 890 bytes per token. These memory figures are not throughput or latency measurements of Lyceum's endpoint.

Smaller KV memory can help an inference engine hold longer sequences, but it does not guarantee a commercial prompt-cache hit or a particular serving speed. DeepSeek's model card specifies up to one million tokens of context; check the hosted route's accepted limits separately. For the earlier model, see the DeepSeek-V4-Flash page.

Pricing: $0.50 input, $0.13 cached, $1.50 output

Token tierPrice per 1M tokens
Input$0.50
Cached input$0.13
Output$1.50

Prices are in USD per 1M tokens, checked against Lyceum's live price book on 17 September 2026. The $0.13 cached-input rate is 74% below $0.50, or 26% of the uncached rate. Only input tokens billed as cache hits receive that discount; repetition alone does not guarantee it. Uncached input and output keep their own rates.

An agent that resends its system prompt, tool definitions and history may benefit when those tokens qualify as cache hits. Total savings depend on the hit rate, uncached input and output volume. The 74% input-tier discount is not a 74% reduction in the whole request or agent loop. For a broader comparison, see Finding the Cheapest Open Model That Clears Your Quality Bar.

Calling deepseek/deepseek-v4.1-flash through the OpenAI-compatible API

Use a Lyceum API key, the base URL https://api.lyceum.technology/openai/v1 and model ID deepseek/deepseek-v4.1-flash. Set LYCEUM_API_KEY in your shell before running either example. OpenAI SDK compatibility covers the request shape shown here; validate optional parameters, tools and reasoning controls separately. The 2,048-token cap bounds this short example; longer reasoning tasks may need a higher limit.

Python, using the official OpenAI SDK (install it with uv pip install openai):

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.lyceum.technology/openai/v1",
    api_key=os.environ["LYCEUM_API_KEY"],
)

response = client.chat.completions.create(
    model="deepseek/deepseek-v4.1-flash",
    messages=[{
        "role": "user",
        "content": "Summarise in three bullets: 09:00 errors rose; 09:05 rollback started; 09:10 service recovered.",
    }],
    max_tokens=2048,
)
print(response.choices[0].message.content)

The same request with curl, against the chat completions endpoint:

curl --fail-with-body https://api.lyceum.technology/openai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $LYCEUM_API_KEY" \
  -d '{
    "model": "deepseek/deepseek-v4.1-flash",
    "messages": [{
      "role": "user",
      "content": "Summarise in three bullets: 09:00 errors rose; 09:05 rollback started; 09:10 service recovered."
    }],
    "max_tokens": 2048
  }'

DeepSeek reports temperature 1.0 and top_p 0.95 for the evaluations below, with numeric reasoning effort set to 100. The upstream model supports effort values from 1 to 100; that does not establish how Lyceum exposes the control.

V4.1-Flash against V4-Pro and V4-Flash on benchmarks

The scores below are percentages from DeepSeek's instruct-model evaluation, not measurements on Lyceum. DeepSeek reports maximum reasoning effort (100), temperature 1.0 and top_p 0.95. The agent scaffolds differ: Terminal-Bench uses DSH Minimal, DeepSWE uses mini-SWE-agent, and AutomationBench uses its official scaffold. DeepSeek reports a 1M context window for its code-agent evaluations. Results depend on the scaffold, tools and test settings as well as the model.

BenchmarkDeepSeek-V4.1-FlashDeepSeek-V4-ProDeepSeek-V4-Flash
Terminal-Bench 2.1 (Pass@1)90.687.982.7
DeepSWE v1.1 (Resolved)74.262.754.4
AutomationBench (Pass@1)54.843.237.7
HLE with tools (Pass@1)63.960.051.5
GPQA Diamond (Pass@1)90.992.489.9

V4.1-Flash leads the selected agentic rows, while V4-Pro leads GPQA Diamond, 92.4 to 90.9. This is a selection, not the complete benchmark suite: DeepSeek also reports V4-Pro ahead on Humanity's Last Exam without tools. These vendor results are useful screening evidence, not a guarantee for your repository or task distribution.

Moving a V4 Pro workload to V4.1-Flash

V4.1-Flash is worth testing for agent workloads given its lower Lyceum token rates and the vendor results above. Migration is not forced by a confirmed V4-Pro retirement: DeepSeek's current API documentation says V4-Pro will continue after 14 September 2026. Both model IDs also appeared in Lyceum's live callable roster when checked on 17 September 2026.

For an existing Lyceum integration, change deepseek/deepseek-v4-pro to deepseek/deepseek-v4.1-flash and keep the same base URL and valid Lyceum key. Then test supported request options, tool calls, output quality, latency and total token cost on representative traffic before switching production. GPQA is one reminder that a lower price does not make every workload better.

  • Price on Lyceum: $0.50 input, $0.13 cached input, $1.50 output per 1M tokens, no minimum commitment
  • Call it with model string deepseek/deepseek-v4.1-flash on https://api.lyceum.technology/openai/v1
  • DeepSeek reports stronger results on the selected agentic benchmarks; V4-Pro still leads some reasoning tests. Validate the choice on your workload.

If you are weighing self-hosting the MIT-licensed weights against calling the API, read Self-Host vs API: The Token Break-Even for Open Models. To run the hosted model, use Serverless Inference. For Claude Code, follow Lyceum's current setup guide: install and authenticate the Lyceum CLI, start lyceum code, then enter /model deepseek/deepseek-v4.1-flash. This uses a separate Lyceum profile.