Hidden-State Probe Detects Physical Danger in LLM Planning Beyond Text Safety
Researchers demonstrate that physical danger (PD) arising from LLM-generated plans for embodied agents is representationally distinct from content danger (CD) found in text. They introduce PRISM, a logistic probe that achieves 86–88% accuracy on SafeAgentBench and 99.6% on the new PhysicalSafetyBench-1K, with significantly fewer false positives than LLM-based judges. This approach enables detection of unsafe actions that are not flagged by traditional text-level moderation.
Why it matters: As LLMs are increasingly used to control robots and physical systems, this work addresses a critical safety gap by providing a method to detect physically unsafe actions that may appear linguistically safe.
Full story at: arXiv Cryptography and Security ↗