← Back to brief
ResearchOfficialPreprintarXiv Computation and Language

LLMs Underperform Encoder-Only Models in Predicting Test Item Difficulty, Raising Concerns for Automated Assessment

A new arXiv preprint reports that large language models (LLMs), including GPT-4.1 and GPT-5.4, are less accurate than encoder-only models like ConvBERT at predicting item difficulty in a large-scale Reading and Writing test. The best-performing LLM (zero-shot GPT-4.1) achieved a QWK of 0.578, compared to ConvBERT's 0.625, and all LLMs struggled particularly with labeling hard items. The study cautions that LLMs may not reliably generate test items with targeted difficulty levels.

Why it matters: This finding highlights a potential limitation of LLMs for educational assessment, where precise control over item difficulty is essential for fair and effective testing.

Full story at: arXiv Computation and Language