[NeurIPS 2025] Official code for "Tropical Attention: Neural Algorithmic Reasoning for Combinatorial Algorithms"
-
Updated
May 21, 2026 - Python
[NeurIPS 2025] Official code for "Tropical Attention: Neural Algorithmic Reasoning for Combinatorial Algorithms"
This work provides extensive empirical results on training LMs to count. We find that while traditional RNNs trivially achieve inductive counting, Transformers have to rely on positional embeddings to count out-of-domain. Modern RNNs (e.g. rwkv, mamba) also largely underperform traditional RNNs in generalizing counting inductively.
Mamba Modulation: On the Length Generalization of Mamba Models
Modular Arithmetic Challenge. Neural induction of exact (a x b) mod p through abacus embeddings, algorithmic scratchpads and grokking, for the SAIR Foundation competition.
Kinship reasoning without a neural network: a sound, set-valued model that derives its rules from opaque symbols. CLUTRR, all six datasets.
Code and experiments for studying length generalization in recurrent Transformers on permutation state-tracking tasks.
Dilution attention: a parameter-free causal attention operator that divides each query's bid by the demand its key has already absorbed. Fused Triton kernels, and a controlled study where 209M models with no position encoding retrieve at 8x their training context.
A modular, config-driven PyTorch framework for conducting empirical positional encoding ablation experiments on decoder-only transformers (~51M parameters) using Google Colab T4 GPUs
An original convergence-reasoning LLM architecture: iterated reasoning emerges from end-answer labels alone and generalizes to 2.3× training depth. PyTorch, ~333K params, fully reproducible.
To associate your repository with the length-generalization topic, visit your repo's landing page and select "manage topics."