← Back to brief
ResearchOfficialPreprintarXiv AI/ML

Study Finds LLMs Struggle to Discern Reliable Sources and Truthfulness, Even as Models Grow

Jul 23, 2026

A new arXiv preprint introduces Learn2Discern (L2D), a benchmark designed to test large language models' (LLMs) ability to weigh information from external sources. Evaluating 13 models across nearly 670,000 trials, the study finds that LLMs perform near chance at distinguishing reliable sources and updating beliefs toward the truth. While newer and larger models show some improvement in truth discernment, they do not improve at recognizing source reliability, highlighting a persistent limitation.

Why it matters: This finding raises concerns about the reliability of LLMs as they are increasingly used to access and evaluate information online.

Full story at: arXiv AI/ML

More coverage