Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning
A study of 4,181 math problems finds that hierarchical multi-agent systems with dedicated reviewer roles do not always outperform simpler broadcast-style peer discussion, especially on harder problems. The performance gap is not due to reviewer precision—PER's reviewer is more precise (0.861 vs. 0.644)—but because critiques are less likely to be acted upon. Forcing explicit acknowledgment of critiques lowers accuracy, while embedding reviewer guidance in the solver's context helps but does not close the gap.
Why it matters: This challenges the assumption that adding a reviewer role inherently improves multi-agent reasoning, showing that the uptake of critiques is a distinct bottleneck from error detection.
Full story at: arXiv AI/ML ↗