Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks, and OpenAI’s accidental AI hacker
Epoch and METR have released MirrorCode, a benchmark designed to evaluate AI systems on long-horizon programming tasks. Current AI systems are still unable to solve the most challenging tasks. The newsletter also discusses the bitter lesson for robotics and an incident involving OpenAI’s accidental AI hacker.
Why it matters: MirrorCode offers a new way to rigorously assess AI's capabilities on extended programming tasks, revealing current limitations and informing future research directions.
Full story at: Import AI — Jack Clark ↗