← Back to brief
ResearchOfficialPreprintarXiv Cryptography and Security

Defender-Centric Evaluation of Jailbreak Attacks via Shapley-Based Framework

Researchers introduce A-MESS, a framework that evaluates jailbreak attacks on large language models (LLMs) based on their utility for improving model safety when used as red-teaming data, rather than solely on attack success rate (ASR). Using Shapley values, the method attributes marginal utility to individual attacks and selects compact, effective subsets for safety training. Experiments show that ASR rankings are only weakly aligned with defender-centric utility, and that optimizing attack subsets directly for safety yields better results than traditional attacker-centric approaches.

Why it matters: This work reframes the evaluation of jailbreak attacks to focus on their value for enhancing LLM safety, potentially leading to more effective red-teaming and safety training strategies.

Full story at: arXiv Cryptography and Security