← Back to brief
ResearchOfficialPreprintarXiv Software Engineering

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

This preprint introduces EvalSafetyGap, a conceptual framework for analyzing failures in large language model (LLM) evaluation and safety under optimization pressure. The authors combine a hybrid survey of eight evidence streams from 2018-2026 with a structured audit of ten LLMs, finding that the link between model capability and adversarial robustness is statistically indeterminate. The study also finds that the observed safety gap between open and closed models is modest and primarily influenced by governance and disclosure practices rather than behavioral robustness.

Why it matters: EvalSafetyGap provides a shared vocabulary and evidence map to improve dynamic evaluation, transparent reporting, and auditable alignment practices in LLM safety research.

Full story at: arXiv Software Engineering

More coverage