What Happened
Researchers published Gated DeltaNet-2 on May 21, 2026, alongside its official implementation. The architecture separates decisions about erasing existing associations and writing new information into recurrent memory.
Why It Matters
Linear-attention models compress history into a fixed-size state, making selective memory updates especially important. Gated DeltaNet-2 develops this alternative to a growing transformer KV cache while retaining efficient parallel training methods.
Technical Details
The update combines channel-wise decay with independent channel-wise erase and write gates, generalizing earlier Gated DeltaNet and Kimi Delta Attention formulations. The authors compare recurrent and hybrid models at 1.3 billion parameters, trained on 100 billion FineWeb-Edu tokens with matched state sizes. They report improvements in language modeling and retrieval within that experimental setting. The implementation uses NVIDIA's noncommercial source-code license, so public availability should not be confused with unrestricted open-source licensing or permission for commercial use.