Apple Introduces VICIS: Benchmarking Visual Concept Inference from Image Sets
Apple Machine Learning Research has introduced Visual Concept Inference from Sets (VICIS), a new task designed to evaluate whether vision-language models (VLMs) can infer shared concepts from small sets of example images and apply them to new queries. The research finds that current state-of-the-art VLMs perform poorly on this benchmark, revealing a significant limitation in their visual reasoning abilities.
Why it matters: This benchmark highlights a key gap in vision-language models' ability to learn and generalize visual concepts from limited visual context, which is important for advancing few-shot learning in AI.
Full story at: Apple Machine Learning Research ↗