Loopie: Looped Transformers Outperform Vanilla Baselines with Same Compute Budget
Researchers introduce Loopie, a looped Transformer architecture that outperforms vanilla Transformer baselines when trained with the same compute budget. The Loopie series includes 20B-parameter (2B active) and 6B-parameter (0.6B active) Mixture-of-Experts models. Extensive ablation studies show Loopie achieves superior performance, including gold-medal results at the 2025 IMO and IPhO benchmarks without external tools.
Why it matters: Loopie demonstrates that looped Transformer architectures can surpass traditional Transformers under equal compute constraints, suggesting a new direction for efficient model scaling.
Full story at: arXiv Computation and Language ↗