Training-Free Frame Selection for Long Videos Using Attention-Based MLLM Selectors
Researchers introduce DAFS, a training-free method for selecting frames in long videos by leveraging cross-modal attention from multimodal large language models (MLLMs). DAFS identifies query-relevant frames without requiring autoregressive generation or additional training, and formulates frame selection as a discrete optimization problem. The approach improves performance over uniform sampling by up to 6.4 points on Video-MME under a 32-frame budget and generalizes across different model backbones and tasks.
Why it matters: This work offers a practical advance for efficient, query-aware frame selection in long-video understanding, potentially broadening the applicability of MLLMs to real-world video analysis tasks.
Full story at: arXiv Computer Vision ↗