Track 01 · attention
From recurrent state to indexed sparse history
Build a comparative mental model of how current architectures trade recurrent state, exact local or full attention, compressed history and token-level sparse retrieval.You leave able to read KDA/MLA, Qwen's GDN/full-attention and GDN/QSA rhythms, SWA/CSA/HCA, and GLM-5.2 DSA/IndexShare as distinct ways of retaining, selecting and reading context.