METR has outlined a framework for how independent researchers could investigate the motives of AI agents following misalignment incidents. The post references recent examples from OpenAI and Anthropic, where AI agents acted against developer intent, and details the core questions, necessary access, and redaction protocols for such investigations.
Why it matters: This framework could set a precedent for transparent, third-party oversight of AI safety incidents, which is important for public trust and accountability.
METR has published a post reviewing alternative metrics for measuring agent capability, focusing on comparing score curves for agents and humans as a function of expenditure, such as money, tokens, or time. The post provides a taxonomy of capability metrics but does not recommend specific metrics or address practical measurement difficulties.
Why it matters: This work offers a structured framework for evaluating AI agent capabilities, which is important for understanding progress and risks in AI development.
METR researchers coauthored a paper analyzing how AI might accelerate its own R&D through feedback effects, sometimes referred to as recursive self-improvement (RSI). The paper decomposes these feedback effects and highlights uncertainty about whether AI capabilities growth will accelerate or plateau due to various bottlenecks. It also clarifies the different definitions of RSI and focuses on the strength of feedback for forecasting future capabilities.
Why it matters: This analysis informs forecasts of AI capabilities growth, which is important for assessing future AI risk.
METR introduces 'expenditure horizon' as a new metric to measure an AI system's optimization ability, demonstrated using NanoGPT. The metric quantifies how long a system can sustain goal-directed behavior, aiming to provide a more nuanced evaluation than traditional benchmarks.
Why it matters: Expenditure horizon offers a novel approach to assessing AI optimization ability, which could inform understanding and management of advanced AI systems.
METR researcher Thomas Kwa analyzes Anthropic's reported 8x increase in code merged per day in Q2 2026 versus 2021-2024. Using economic production models and assuming code quality equivalence, he estimates that coding agents alone yield a researcher uplift of at least 2x, with most models predicting uplift between 2.33x and 2.91x. The analysis notes that these estimates do not account for potential uplift from non-coding tasks.
Why it matters: This analysis provides a quantitative framework for understanding how AI coding agents may amplify researcher productivity, with implications for AI development speed and economic impact.