Salience Induction: A New Attack Surface for Multi-Hop RAG Agents
Researchers have identified a novel attack surface in agentic retrieval-augmented generation (RAG) systems, termed 'salience induction.' Unlike content poisoning or prompt injection, this attack manipulates the prominence and framing of truthful information to mislead multi-hop reasoning agents. The team introduces a defense method called Salience Normalization, which significantly reduces attack success rates from 83.3% to 15.3% in their experiments across multiple model families and agent architectures.
Why it matters: This work exposes a critical vulnerability in RAG systems, showing that even truthful content can be manipulated to mislead AI agents, and demonstrates an effective defense.
Full story at: arXiv Cryptography and Security ↗