Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning
Researchers have identified a critical failure mode in long-context large language models (LLMs) called repetitive copying, where models copy input text into their reasoning traces instead of engaging in productive problem-solving. They introduce GEAR, a reward shaping method that encourages grounding in key evidence and penalizes copying from irrelevant context, leading to consistent improvements of up to +4.6 average points over standard reinforcement learning approaches across multiple benchmarks and model scales.
Why it matters: This work highlights and addresses a pervasive limitation in long-context LLMs, demonstrating that improved evidence grounding can significantly enhance reasoning performance and reduce unproductive copying.
Full story at: arXiv Computation and Language ↗