RIMS: Preference Optimization via Smoothed Multi-pair Aggregation for Small-Scale LLM Retrieval-Augmented Generation
Researchers introduce RIMS, a three-stage preference optimization framework designed to enhance small language models in retrieval-augmented generation tasks. RIMS generates synthetic chain-of-thought preference data, employs a differentiable soft aggregation mechanism to better utilize preference signals, and applies preference optimization to improve robustness against noisy evidence. Experiments on four multi-hop question answering benchmarks demonstrate that RIMS consistently outperforms state-of-the-art baselines in both Exact Match and F1 scores under noisy retrieval conditions.
Why it matters: This work advances the reliability and effectiveness of small language models for retrieval-augmented generation, particularly in resource-constrained environments.
Full story at: arXiv Computation and Language ↗