Hide and Seek in Embedding Space: Geometry-based Steganography and Detection in Large Language Models
Researchers formalize the concept of payload recoverability in LLM steganography and introduce new low-recoverability schemes using embedding-space mappings. Their experiments show that while these schemes reduce the ease of secret recovery, mechanistic interpretability methods—specifically linear probes on later-layer activations—can still detect steganographic content with up to 33% higher accuracy in fine-tuned models compared to base models.
Why it matters: This work highlights both the evolving threat of covert communication via LLMs and the potential for interpretability-based defenses, signaling a new dynamic in model security.
Full story at: arXiv Cryptography and Security ↗