GAMUT: A Benchmark for Factual Completeness in Long-Form Generation
Researchers have introduced GAMUT, a benchmark designed to evaluate factual completeness in long-form AI generation. GAMUT employs a two-level meta-rubric framework to assess whether AI-generated responses include all necessary information, rather than just avoiding factual errors. The benchmark features 1,813 questions across 10 domains, and the best-performing model (Gemini 3.1 Pro) achieved a score of 58.7%.
Why it matters: This work provides a structured and rigorous method to assess whether AI-generated long-form content is fully informative, addressing a key gap in current evaluation practices.
Full story at: arXiv Computation and Language ↗