Benchmark Finds AI Image Detectors Unreliable in Safety-Critical Scenarios
A new arXiv preprint introduces SafeIMG, a benchmark designed to test AI-generated image detectors in 12 scenarios relevant to public and individual safety. The study finds that leading vision-language models and specialized detectors perform far below human accuracy, with the best model detecting only about half of synthetic images and providing limited explanations for anomalies. Detection and explanation performance drops further for commonsense and physical inconsistencies, and after image degradation.
Why it matters: The results highlight significant limitations in current AI image detection tools, raising concerns about their reliability in high-stakes contexts where visual authenticity is crucial.
Full story at: arXiv Computer Vision ↗