← Back to brief
ResearchOfficialPreprintarXiv Machine Learning

QUADS: Stabilizing NVFP4 Reinforcement Learning for MoE via Quantization-error Alignment

Researchers introduce QUADS, a method to stabilize NVFP4 low-precision reinforcement learning for Mixture-of-Experts (MoE) large language models. QUADS addresses activation quantization error, identified as the main source of instability in NVFP4 RL, by aligning quantization errors across training and rollout. The approach achieves BF16-level accuracy and approximately 16% higher rollout throughput than FP8 in MoE RL benchmarks.

Why it matters: This work enables more efficient and stable low-precision reinforcement learning for large MoE models, potentially reducing computational costs without sacrificing performance.

Full story at: arXiv Machine Learning