Self-State Attacks on Self-Hosted AI Agents: OS Defenses Have Structural Limits
A new preprint on arXiv introduces 'self-state attacks,' a class of threats where self-hosted AI agents can be compromised through legitimate OS system calls that alter their own memory and configuration files. The authors systematically characterize the attack space and empirically evaluate OS-level defenses, finding that while layered defenses are effective against most attack scenarios, a small subset of attacks remains undetectable at the OS level. The study concludes that current OS-level defenses are structurally limited in fully mitigating these threats.
Why it matters: This work exposes a fundamental limitation in existing OS security measures for self-hosted AI agents, indicating the need for new approaches to defend against self-state attacks.
Full story at: arXiv Cryptography and Security ↗