← Back to brief
ResearchOfficialPreprintarXiv Machine Learning

HARP: Training-Free Interpretability with a Tool-Using Agent

A new preprint introduces HARP, a training-free interpretability method that equips a large language model (LLM) agent with a vector database of neural activations and tools for manipulating them. HARP enables the agent to iteratively retrieve, hypothesize, and validate concepts in neural networks without additional training. The method reportedly outperforms training-based approaches like sparse autoencoders (SAEs) and activation oracles on tasks such as concept discovery, model steering, and secret elicitation, while being more cost-effective and flexible.

Why it matters: HARP challenges the assumption that training-based interpretability methods extract fundamentally deeper insights, motivating the development of new benchmarks for interpretability.

Full story at: arXiv Machine Learning