VIABench: A Video Benchmark for Evaluating MLLMs in Visually Impaired Assistance
Researchers have introduced VIABench, a video benchmark specifically designed to evaluate Multimodal Large Language Models (MLLMs) in the context of assisting visually impaired individuals. VIABench uses first-person videos from blind individuals and defines three core tasks: Proactive Reminder, Visual Question Answering, and Vision-Guided Interaction. Experimental results indicate that current MLLMs face significant challenges, particularly in proactive anticipation and real-time responsiveness.
Why it matters: VIABench highlights critical gaps in current MLLMs for real-world blind assistance, providing a new resource to drive research toward more effective navigation and interaction support for visually impaired individuals.
Full story at: arXiv Computer Vision ↗