← Back to brief
ResearchOfficialPreprintarXiv Machine Learning

Normalized Rewards for Preference Optimization Improve LLM Alignment and Generalization

A new preprint proposes adding a regularization term to Direct Alignment Algorithms (DAAs) such as DPO and SimPO to maintain length-normalized probabilities of chosen and rejected responses, addressing over-optimization issues. The method leads to improved trade-offs between generation quality and general benchmark performance, with reported gains including over 20% relative increase in AlpacaEval2 scores and over 9% improvement on general benchmarks for Llama-3.1-8B-Instruct. The regularization also reduces undesirable likelihood shifts, particularly for outlier tokens.

Why it matters: This work offers a practical solution to a known limitation in preference optimization for LLMs, enhancing both alignment and general capabilities with a simple regularization technique.

Full story at: arXiv Machine Learning