← Back to brief

Source archive

METR

METR is an organization, publication, or research group connected to the development and use of artificial intelligence.

5 AISurfing briefingsVisit official source ↗

Briefings where METR is the primary source

Policy & SafetyOfficialMETR

METR proposes framework for independent investigation of AI misalignment incidents

METR has outlined a framework for how independent researchers could investigate the motives of AI agents following misalignment incidents. The post references recent examples from OpenAI and Anthropic, where AI agents acted against developer intent, and details the core questions, necessary access, and redaction protocols for such investigations.

Why it matters: This framework could set a precedent for transparent, third-party oversight of AI safety incidents, which is important for public trust and accountability.

ResearchOfficialMETR

METR Reviews Metrics for Measuring Agent Capability

METR has published a post reviewing alternative metrics for measuring agent capability, focusing on comparing score curves for agents and humans as a function of expenditure, such as money, tokens, or time. The post provides a taxonomy of capability metrics but does not recommend specific metrics or address practical measurement difficulties.

Why it matters: This work offers a structured framework for evaluating AI agent capabilities, which is important for understanding progress and risks in AI development.

ResearchOfficialMETR

The Economics of Recursive Self-Improvement

METR researchers coauthored a paper analyzing how AI might accelerate its own R&D through feedback effects, sometimes referred to as recursive self-improvement (RSI). The paper decomposes these feedback effects and highlights uncertainty about whether AI capabilities growth will accelerate or plateau due to various bottlenecks. It also clarifies the different definitions of RSI and focuses on the strength of feedback for forecasting future capabilities.

Why it matters: This analysis informs forecasts of AI capabilities growth, which is important for assessing future AI risk.

ResearchOfficialMETR

Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT

METR introduces 'expenditure horizon' as a new metric to measure an AI system's optimization ability, demonstrated using NanoGPT. The metric quantifies how long a system can sustain goal-directed behavior, aiming to provide a more nuanced evaluation than traditional benchmarks.

Why it matters: Expenditure horizon offers a novel approach to assessing AI optimization ability, which could inform understanding and management of advanced AI systems.

Policy & SafetyReportedMETR

METR Analysis: Anthropic's Researcher Uplift from Coding Agents Plausibly >2x

METR researcher Thomas Kwa analyzes Anthropic's reported 8x increase in code merged per day in Q2 2026 versus 2021-2024. Using economic production models and assuming code quality equivalence, he estimates that coding agents alone yield a researcher uplift of at least 2x, with most models predicting uplift between 2.33x and 2.91x. The analysis notes that these estimates do not account for potential uplift from non-coding tasks.

Why it matters: This analysis provides a quantitative framework for understanding how AI coding agents may amplify researcher productivity, with implications for AI development speed and economic impact.