← Back to brief
ResearchOfficialPreprintarXiv Computer Vision

MoD-VLLM: Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding

A new framework called MoD-VLLM is proposed for multi-event long video understanding. The approach iteratively localizes question-relevant video segments using a grounding module and a reflection module with dynamic granularity encoding. It employs a reinforcement learning strategy to jointly optimize grounding policies and visual representations. MoD-VLLM demonstrates significant improvements over state-of-the-art baselines on several benchmarks, including the newly introduced MEventBench.

Why it matters: This work introduces a modular and self-reflective approach that addresses the challenge of understanding multiple events in long videos, a key limitation of current Video LLMs.

Full story at: arXiv Computer Vision