Few-Shot Prompting Effects Vary Widely Across LLMs: Systematic Study Reveals Non-Monotonic and Unpredictable Patterns
A new arXiv preprint presents a controlled study of five large language models (LLMs) across six few-shot prompting configurations on the AG News benchmark. The authors find that the effect of adding more examples (shots) is highly variable: some models show no improvement, some recover from poor zero-shot performance, others degrade with more examples, and one exhibits a U-shaped performance curve. The study also uncovers a parsing artifact that significantly distorted results for one model, highlighting the importance of robust evaluation methods.
Why it matters: These findings challenge the common assumption that more prompt examples always help, revealing that few-shot prompting effects are complex and not reliably predicted by model size or architecture.
Full story at: arXiv Information Retrieval ↗