Frontier Model Atlas

DeepSeek · DeepSeek V4

DeepSeek V4 Flash

A 284B / 13B-active MoE model built for million-token context with hybrid CSA/HCA attention, mHC residual streams and one MTP layer.

Total284BActive13BLayers43Context1M
Checkpoint architecture · source-backed layout
Model anatomyToken path · block microscope · cost surface
Architecture signature2 SWA · 21 CSA · 20 HCA
SWA2 · 4.7%CSA21 · 48.8%HCA20 · 46.5%
Layer blueprintAttention · FFN · residual, aligned by transformer block43 layers
Layer mix
2× SWA · 21× CSA · 20× HCA
FFN schedule
3 hash-routed MoE · 40 learned-routing MoE
Experts / token
6 / 256 routed + 1 shared
Residual topology
4-stream mHC
Active parameter share
4.6% of checkpoint

Each blueprint column is one transformer block. Hover or keyboard-focus a linked cell to read attention, feed-forward and residual structure as one aligned layer; linked cells open the corresponding walkthrough or inspector. DSA cells additionally preserve whether the layer runs a full indexer or reuses a shared IndexShare selection. QSA cells preserve their per-layer sparse retrieval identity, while PLE is marked on the decoder layer receiving N-gram features. MoE cells preserve any layer-level routing transition encoded by the checkpoint manifest. Structural encoding only — it does not imply benchmark quality, throughput or FLOPs.