CaRE: Compute-Aware Evaluation Protocol Reveals Flaws in Masked Diffusion Language Model Benchmarks
A new evaluation protocol, CaRE, standardizes compute-aware comparisons for masked diffusion language models (MDLMs) by controlling for function evaluations, reporting multiple metrics, and explicitly managing stochasticity. The study finds that temperature settings account for most of the variance in a key evaluation metric (MAUVE), and that previously published rankings of remasking strategies can reverse when compute is matched. This suggests that many prior claims about MDLM improvements may be confounded by inconsistent evaluation practices.
Why it matters: The work highlights that widely used evaluation methods for MDLMs may systematically misattribute algorithmic gains, underscoring the need for standardized, reproducible benchmarks in this fast-moving area.
Full story at: arXiv AI/ML ↗