← Back to brief
ResearchOfficialPreprintarXiv Cryptography and Security

HALLMARK Benchmark Identifies False-Positive Rate as Main Challenge for LLM Citation Verification

A new preprint introduces HALLMARK, a benchmark designed to evaluate large language model (LLM) citation verifiers using 2,526 BibTeX entries and 14 types of citation hallucinations. The study finds that the false-positive rate, rather than recall, is the primary factor limiting the practical deployment of citation verification tools, with most LLMs tending to over-flag citations, especially for papers published after their training cutoff.

Why it matters: This benchmark provides a standardized way to diagnose and improve LLM citation verifiers, addressing the risk of fabricated references in AI-generated academic writing.

Full story at: arXiv Cryptography and Security