← Back to brief
ResearchOfficialPreprintarXiv Computation and Language

Trace-Based On-Policy Distillation Boosts Reasoning in Diffusion LLMs

Researchers introduce Trace-Based On-Policy Distillation (TOPD), a teacher-supervised framework for transferring reasoning abilities to masked diffusion language models without the need for reward estimation. TOPD enables the SDAR-4B-Chat model to achieve comparable accuracy to RL-trained models on the MATH500 benchmark, while requiring four times fewer rollout rounds and an estimated 96-fold compute speedup.

Why it matters: This approach offers a more efficient method for reasoning-oriented post-training of diffusion LLMs, potentially reducing computational costs compared to reinforcement learning-based methods.

Full story at: arXiv Computation and Language