RL with Verifiable Rewards Can Harm High-Budget Coverage: Diagnosing Pass@k Inversion
A new arXiv preprint reports that reinforcement learning with verifiable rewards (RLVR) can paradoxically reduce a model’s ability to solve problems when allowed multiple attempts, despite improving single-sample accuracy. This 'pass@k inversion' occurs because RLVR may eliminate rare correct solutions that only appear with repeated sampling, especially on boundary prompts. The authors introduce Per-Problem Base Anchoring (PBA) as a proof-of-concept method to mitigate this effect by preserving rare correct trajectories.
Why it matters: The findings highlight a fundamental risk in verifier-guided RL training that could impact the reliability of AI systems in settings where repeated attempts are important, such as vision-language agents or mathematical reasoning tasks.
Full story at: arXiv Machine Learning ↗