← Back to brief
ResearchOfficialPreprintarXiv Computation and Language

OmniaBench: New Benchmark Evaluates General AI Agents Across Diverse Scenarios

Researchers have introduced OmniaBench, a benchmark designed to evaluate general AI agents across a wide range of scenarios with explicit state spaces. The benchmark spans 90 level-1 and 354 level-2 domains, covering consumer, business, and enterprise contexts, and includes 1,431 tasks. Leading AI models such as Claude-Sonnet-5 and GPT-5.6-Sol achieved only 58.54% and 57.14% Overall Pass@1 scores, highlighting ongoing challenges in planning and adaptive correction.

Why it matters: OmniaBench offers a comprehensive and diagnostic tool for systematically assessing the capability boundaries of general AI agents across heterogeneous real-world applications.

Full story at: arXiv Computation and Language