← Back to brief
ResearchOfficialPreprintarXiv Software Engineering

Teach it to stop, not just to click: Reproducibility and repair in agentic computer-use RL

A preprint investigates reproducibility in reinforcement learning for agentic computer-use, focusing on a 35B-parameter agent. The study finds that single-run evaluations are unreliable due to high variance from data sampling and nondeterminism, with evaluation variance itself being negligible. Repairability is shown to be two-tiered: fixed-token interventions are reliably effective, while open-ended corrections are only partially successful. The authors also release a library to facilitate routine multi-seed evaluation reporting.

Why it matters: This work exposes critical reproducibility challenges in agentic RL and introduces practical tools and analysis for more reliable evaluation and repair of computer-use agents.

Full story at: arXiv Software Engineering