← Back to brief
ResearchOfficialPreprintarXiv Computer Vision

CRISP: Pre-LLM Text-Driven Visual Token Pruning for Efficient LVLM Inference

Researchers introduce CRISP, a two-stage framework for pruning visual tokens before input to large vision-language models (LVLMs), guided by text to retain instruction-relevant evidence and scene context. Experiments on LLaVA-1.5 and LLaVA-NeXT show that CRISP can maintain up to 99.5% of original accuracy while reducing inference cost and latency by more than 2x. This approach addresses the inefficiency of processing large numbers of visual tokens in LVLMs, particularly for resource-constrained environments.

Why it matters: CRISP enables more efficient deployment of LVLMs by significantly reducing inference overhead without major accuracy loss, making advanced multimodal AI more accessible in limited-resource settings.

Full story at: arXiv Computer Vision