← Back to brief
ResearchOfficialPreprintarXiv AI/ML

FindStatBench: New Benchmark Reveals LLM Strengths and Weaknesses in Combinatorial Code Synthesis

FindStatBench is a new execution benchmark that evaluates large language models (LLMs) on combinatorial code synthesis, featuring 2,329 tasks and 5.52 million hidden instances. The benchmark shows that top open- and closed-source LLMs achieve nearly identical accuracy, but providing examples can sometimes reduce performance, and long prompts can sharply decrease accuracy. The results highlight that exact symbolic rule induction remains a significant challenge for current LLMs.

Why it matters: FindStatBench exposes fundamental limitations in LLMs' ability to perform symbolic reasoning and code synthesis, informing future research directions.

Full story at: arXiv AI/ML