Study Finds LLM 'Reflection' Fails to Improve Answers Like Human Revision
A new arXiv preprint introduces the Human-LLM Reflection Framework (HRF) to directly compare how humans and large language models (LLMs) revise their answers. The study finds that, unlike humans, LLMs gain little or even negative information from self-reflection: on objective tasks, LLMs' revisions are essentially neutral, while on subjective tasks, they often move further from the correct answer. The failure is traced to the revision process itself, not the quality of initial responses, suggesting that LLM 'reflection' is more akin to re-generating text than genuine error correction.
Why it matters: The findings challenge the effectiveness of prompting LLMs to 'reflect' and suggest that current self-revision methods may not improve model reasoning as previously hoped.
Full story at: arXiv Machine Learning ↗