Kernelized Linear Attention Breaks Capacity Wall with Symmetric Cones
Researchers introduce KATA, a linear attention framework derived from first principles using self-dual homogeneous cones to ensure nonnegative attention weights. KATA achieves up to 11x the throughput of FlashAttention-2 at 131k tokens and maintains near-perfect associative recall. On long-range tasks, KATA variants outperform Gated DeltaNet and retain high MQAR (0.985) at 16x out-of-distribution sequence lengths.
Why it matters: KATA provides a principled approach to overcoming the capacity-interference tradeoff in linear attention, enabling efficient and accurate long-context inference.
Full story at: arXiv Statistical ML ↗