DeepSeek · DeepSeek V4
DeepSeek V4 Flash
A 284B / 13B-active MoE model built for million-token context with hybrid CSA/HCA attention, mHC residual streams and one MTP layer.
DeepSeek · DeepSeek V4
A 284B / 13B-active MoE model built for million-token context with hybrid CSA/HCA attention, mHC residual streams and one MTP layer.
Each blueprint column is one transformer block. Hover or keyboard-focus a linked cell to read attention, feed-forward and residual structure as one aligned layer; linked cells open the corresponding walkthrough or inspector. DSA cells additionally preserve whether the layer runs a full indexer or reuses a shared IndexShare selection. QSA cells preserve their per-layer sparse retrieval identity, while PLE is marked on the decoder layer receiving N-gram features. MoE cells preserve any layer-level routing transition encoded by the checkpoint manifest. Structural encoding only — it does not imply benchmark quality, throughput or FLOPs.