Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control
A new preprint explores how mechanistic interpretability can reveal low-dimensional, robustness-critical features in World Action Models (WAMs), enabling targeted interventions. The authors introduce contrastive activation directions for training-free steering and propose the World-Action Linear Quadratic Regulator (WA-LQR), a feedback control method that leverages local linearity in WAM activation dynamics. Experiments show that these approaches improve robustness to various perturbations in Cosmos-Policy and DiT4DiT models, outperforming baseline methods.
Why it matters: This work offers a novel, training-free approach to enhance the robustness of world action models, which could improve reliability in robotics and autonomous systems.
Full story at: arXiv Robotics ↗