Speed Always Wins: A Survey on Efficient Architectures for Large Language Models
-
Updated
Nov 11, 2025
Speed Always Wins: A Survey on Efficient Architectures for Large Language Models
[NeurIPS 2024] Official code of ”LION: Linear Group RNN for 3D Object Detection in Point Clouds“
101M parameter linear recurrence and attention model: GDN-2 + GQA with 2-pass block recycling
Saryu: a recurrent language model whose state is moved by input-dependent Householder reflections. Open architecture research from India — constant memory per token, an exact parallel kernel, and two technical reports: the architecture, and what the same transport does when it fails (it collapses onto exact group quotients).
Triton kernels for linear RNN (currently recurrentgemma)
Sparse autoencoder for recurrentgemma
minGRU (Feng et al.) in PyTorch, plus a measured ladder of state-tracking recoveries — signed, rotation, Givens, and delta (DeltaNet-style Householder products) mixers with time decay, all under one parallel scan — with fused Triton GPU kernels as an optional backend
To associate your repository with the linear-rnn topic, visit your repo's landing page and select "manage topics."