METR Reviews Metrics for Measuring Agent Capability
METR has published a post reviewing alternative metrics for measuring agent capability, focusing on comparing score curves for agents and humans as a function of expenditure, such as money, tokens, or time. The post provides a taxonomy of capability metrics but does not recommend specific metrics or address practical measurement difficulties.
Why it matters: This work offers a structured framework for evaluating AI agent capabilities, which is important for understanding progress and risks in AI development.
Full story at: METR ↗