Causal Audit Finds Attention-Based Memory Eviction Unreliable in Multimodal Assistants
A new arXiv preprint introduces the Causal Visual Memory Audit (CVMA), a framework for testing when visual information can be safely removed from the memory of multimodal AI assistants. The study finds that current attention-based strategies for evicting visual memory can perform worse than random at retaining information needed for future dialog turns. The results suggest that safe forgetting depends on whether visual facts will be needed again or have been explicitly verbalized, rather than on current attention scores.
Why it matters: This work highlights a key limitation in how multimodal AI assistants manage visual memory, raising concerns about their reliability in long, stateful interactions.
Full story at: arXiv Computer Vision ↗