MissingBench-Verified: VLMs Consistently Fail to Detect Missing Object Parts, Even with Tool Assistance
A new benchmark, MissingBench-Verified, demonstrates that leading vision-language models (VLMs) consistently fail to recognize when an essential part of an object is missing from an image. This failure persists even when external tool evidence, such as image processing outputs, contradicts the model's initial perception. The study finds that current mitigation strategies—including tool-assisted verification, autonomous visual reasoning, and fine-tuning—offer negligible improvement, indicating that this limitation cannot be addressed with existing techniques.
Why it matters: This exposes a fundamental limitation in current VLMs for inspection and monitoring applications, suggesting that more substantial architectural or training changes are required.
Full story at: arXiv Computer Vision ↗