← Back to brief
ResearchOfficialPreprintarXiv Machine Learning

RENEW: Using Human Preferences to Repair World Model Exploitation in Offline RL

Researchers introduce RENEW, a method that leverages human preferences over imagined rollouts to address model exploitation in offline reinforcement learning. The approach, formalized as Dynamics Learning from Human Feedback (DLHF), uses human intuition to identify unrealistic model dynamics. RENEW improves the practicality of preference-based supervision by focusing finetuning on regions where the model is most uncertain, thereby enhancing sample efficiency and reducing exploitation.

Why it matters: This work presents a novel approach to mitigating model exploitation in offline RL by directly incorporating human feedback, potentially reducing reliance on costly expert demonstrations.

Full story at: arXiv Machine Learning