Prompt Design at Scale: Format, Instruction Count, and Context Length Shape LLM Adherence and Hallucination
A new preprint presents controlled experiments examining how prompt format, instruction count, and context length affect large language model (LLM) instruction adherence and hallucination. The study finds that perfect response rates drop to zero by 80 instructions across all tested models and formats, and recall accuracy remains high up to 64-128k tokens before degrading sharply, with significant format-dependent differences. Notably, the experiments observe no fabrication (hallucinated facts) but a sharp rise in outright refusal to answer as models approach their context limits.
Why it matters: This work provides rare, systematic evidence on prompt design tradeoffs, clarifying practical limits on instruction count and context length for LLM reliability.
Full story at: arXiv Computation and Language ↗