ZeroSplat: Training-Free 3D Segmentation from Language Queries
Researchers introduce ZeroSplat, a training-free framework for generalized referring segmentation in 3D Gaussian Splatting. ZeroSplat enables segmentation of zero, one, or multiple targets in response to language queries by transferring 2D vision-language model priors into 3D space using multi-view geometric constraints. The method demonstrates significant performance improvements over existing approaches on the newly proposed GR-LERF and GR-ScanNet benchmarks, without requiring per-scene optimization.
Why it matters: ZeroSplat addresses key limitations in language-guided 3D scene understanding, enabling more flexible and efficient segmentation that better reflects the ambiguity of real-world instructions.
Full story at: arXiv Computer Vision ↗