← Back to brief
ResearchOfficialPreprintarXiv Cryptography and Security

IH-Benchmark: Benchmarking LLM Robustness to Conflicting Instruction Hierarchies

A new arXiv preprint introduces IH-Benchmark, a benchmark designed to evaluate how large language models (LLMs) handle conflicting instructions from different hierarchy levels, such as system-user and user-tool conflicts. Testing 37 models, the study finds that strong compliance with system-user instructions does not reliably predict robustness when conflicts arise via tool outputs. The benchmark spans 44 constraint families across domains including health, finance, and coding, and reveals that instruction-hierarchy robustness is a complex, multi-faceted challenge for LLMs.

Why it matters: Understanding and measuring LLM robustness to conflicting instructions is critical for safe and reliable deployment in real-world applications where such conflicts are common.

Full story at: arXiv Cryptography and Security

More coverage