← Back to brief
ResearchOfficialPreprintarXiv Machine Learning

RRPO: Reference-Relative Policy Optimization with Stratified Conditional Rollouts

Researchers introduce Reference-Relative Policy Optimization (RRPO), a method that generalizes Group Relative Policy Optimization (GRPO) by using reference-relative contrastive comparisons instead of correctness-based advantage construction. RRPO employs stratified conditional rollouts to create positive and negative anchor sets, then trains a metric projection head with a set-contrastive objective to define contrastive advantages for policy optimization. The approach is shown to be competitive with verifier-based optimization across tasks such as verifiable reasoning, open-ended generation, and post-supervised fine-tuning settings.

Why it matters: RRPO broadens the applicability of group-relative optimization methods to tasks lacking a single correctness criterion, potentially enabling reinforcement learning in more diverse and open-ended domains.

Full story at: arXiv Machine Learning