Group Entropy-Controlled Policy Optimization (GEPO) Improves LLM Alignment Across Heterogeneous Tasks
Researchers introduce Group Entropy-Controlled Policy Optimization (GEPO), a lightweight extension to GRPO that leverages group-level entropy to adapt advantage signals for reinforcement learning alignment of large language models (LLMs) on heterogeneous tasks. GEPO addresses the challenge of varying exploration needs across diverse task groups and demonstrates consistent outperformance over GRPO and other entropy-controlled methods on 13 benchmarks, including math, physics, code, and instruction following.
Why it matters: GEPO offers a practical advance for aligning LLMs on mixed-task datasets, enabling more balanced and effective training across diverse capabilities.
Full story at: arXiv Computation and Language ↗