Orthogonalized Read Acts as a Removable Training Scaffold in Recurrent Memory Models
A new arXiv preprint reports that orthogonalizing the memory matrix at read time in mLSTM models improves performance on noisy associative recall tasks, not by increasing memory capacity, but by re-conditioning the learning problem during a training plateau. The intervention is shown to be self-consistent, uniform across hyperparameters, and removable—removing it after training leaves standard mLSTMs at full accuracy. The findings suggest that much of the observed improvement is due to changes in trainability rather than architectural memory enhancements.
Why it matters: This challenges common interpretations of memory architecture benchmarks, indicating that some reported gains may reflect training dynamics rather than true advances in model capacity.
Full story at: arXiv Machine Learning ↗