← Back to brief
Policy & SafetyOfficialPreprintarXiv Cryptography and Security

Membership Inference Attacks in the Unseen Class Setting: Limitations and a More Robust Approach

A new arXiv preprint formalizes the 'unseen class' scenario for membership inference attacks (MIAs), where auditors cannot access representative samples of certain content classes—such as harmful material—due to legal or ethical barriers. The study finds that state-of-the-art MIA techniques perform poorly in this setting, while quantile regression-based attacks can achieve up to 11 times the true positive rate of traditional shadow model-based methods. The authors provide both empirical results and theoretical analysis supporting this improvement.

Why it matters: This work highlights a significant limitation in current data auditing tools for AI safety and introduces a more effective method for detecting problematic training data when access to harmful examples is restricted.

Full story at: arXiv Cryptography and Security

More coverage