Frontier Flux Original

Other· 2 min read

DeepSeek ships V4.1-Flash MIT weights and API deepseek-flash

DeepSeek released DeepSeek-V4.1-Flash with MIT weights on Hugging Face and API id deepseek-flash. The lab describes a 552B MoE Causal Encoder-Decoder that activates 8B parameters on prefill and 16B on decode, plus a 890-byte global KV cache per token.

ShareXEmail
Abstract cyan rings and stacked planes on a dark field, AI-generated

Abstract concept art for DeepSeek V4.1-Flash. AI-generated.

Frontier Flux / AI-generated

DeepSeek published DeepSeek-V4.1-Flash on 10 September 2026: MIT-licensed weights on Hugging Face, a technical report, and API access under `deepseek-flash`. The lab frames it as the smallest model in a new architecture family, with native image input and a 1 million token context window.

From the 10 September 2026 announcement: "Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. Introducing the smallest model in our new architecture family, with native visual understanding. Designed for greater capability, faster inference, higher throughput, and scaling to larger models."

The official news post and model card fill in the architecture. DeepSeek says V4.1-Flash is a 552B-parameter mixture-of-experts with a Causal Encoder-Decoder layout: 8B active parameters per token during prefill, 16B during decode. Combined with Compressed Sparse Attention 2 and FP4 main KV caching, the card reports a global KV cache of 890 bytes per token, about one quarter of V4-Flash, and a persistent cache footprint about one eighth of that prior Flash generation.

Key facts

  • Access: API id `deepseek-flash` on DeepSeek's pricing page; weights at `deepseek-ai/DeepSeek-V4.1-Flash` under MIT.
  • Context: 1M tokens; docs list a 384K max output. Vision is on for Flash. V4-Pro-0813 remains text-only until the planned reroute.
  • Retirements (DeepSeek): V4-Flash and V4-Flash-Vision-Exp are retired; `deepseek-v4-flash` and `deepseek-v4-flash-vision-exp` temporarily route to V4.1-Flash. From 04:00 UTC on 14 September 2026, `deepseek-v4-pro` requests route to V4.1-Flash at Flash rates until V4.1-Pro launches.
  • Pricing (docs, per 1M tokens): Flash peak cache-hit $0.006, cache-miss $0.30, output $1.20. Off-peak is half. Peak hours: 01:00-04:00 and 06:00-10:00 UTC, Monday-Friday.
  • Training (card): trained from scratch on 45T multimodal tokens. Reasoning effort is an integer 1-100. DeepSeek-reported max-effort scores include Terminal-Bench 2.1 at 90.6 and DeepSWE v1.1 at 74.2.

How to read this: the hard ship is weights plus API plus a documented CED and KV-cache design. Rank claims versus V4-Pro sit in DeepSeek's tables and announcement copy.

Frontier Flux is independent of DeepSeek. Social posts are the alert; the news page, Hub card, LICENSE file, and API pricing table are the primaries.

ShareXEmail

Free

Get the morning brief

Free daily email of model releases, benchmarks, and lab notes from the last 24 hours across OpenAI, Anthropic, DeepMind, Meta, NVIDIA, and more. Capability-first, not company drama.

Opens Buttondown to finish signup. Confirm by email if asked. No ads. Unsubscribe anytime.

Back to latest · All Originals · Morning brief