More Than Memory: Task-Conditioned Signed FFN Writes in Long-Context Retrieval
A new preprint investigates the role of feedforward network (FFN) writes in large language models (LLMs) during long-context retrieval tasks. The study finds that FFN writes are not only signed and layer-specific, but also conditioned on the retrieval task. Notably, the final FFN layer typically acts as a suppressor, and a diagnostic based on write-gradient alignment can accurately predict when retrieval performance will degrade.
Why it matters: This work introduces a compact diagnostic tool for analyzing and potentially improving LLM performance on long-context retrieval, a key challenge for real-world applications.
Full story at: arXiv Machine Learning ↗