Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning
A new method, Object-Part Hierarchical Reflective Grounding (OP-HRG), is introduced to address the challenge of part-level visual grounding in multimodal large language models. OP-HRG employs a coarse-to-fine reasoning strategy, first localizing the parent object and then the specific part, with a reflective self-check mechanism. Trained with a part-aware reinforcement learning framework, the approach achieves state-of-the-art results on several part grounding benchmarks, outperforming larger existing models.
Why it matters: This work advances fine-grained visual understanding in multimodal models, enabling more accurate part-level grounding for applications such as robotics and image editing.
Full story at: arXiv Computer Vision ↗