Bad Memory: Evaluating Prompt Injection Risks from Memory in Agentic Systems
A new preprint investigates prompt injection vulnerabilities in agentic systems that use persistent memory, focusing on Anthropic Claude Code and OpenAI Codex across four models. The study finds that while it is challenging for attackers to overwrite memory files directly, malicious payloads already present in memory can persist and compromise both current and future sessions. The persistence and effectiveness of these attacks vary depending on the system, model, and adversarial objectives.
Why it matters: This work demonstrates that persistent memory in agentic systems introduces a novel and significant attack vector for prompt injection, underscoring the need for new defenses that secure memory updates without impeding beneficial adaptation.
Full story at: arXiv Cryptography and Security ↗