← Back to brief
ResearchOfficialPreprintarXiv Computation and Language

ToxGate: Conditional Reliability of Toxicity Signals for Multilingual and Code-Mixed Abuse Detection

A new preprint introduces ToxGate, a trust-fusion head that conditions external toxicity signals on encoder representations to enhance abuse detection in multilingual and code-mixed text. Evaluated across three short-text abuse datasets and multiple transformer encoders, ToxGate outperforms plain encoders in 10 of 12 in-domain and 7 of 8 transfer settings, with the most significant improvements in high-risk moderation slices such as explicit slurs and violent threats. The study demonstrates that treating external toxicity tools as conditional evidence, rather than fixed features, leads to more reliable moderation.

Why it matters: This work provides a practical advance for content moderation by improving the reliability of abuse detection in challenging multilingual and code-mixed scenarios.

Full story at: arXiv Computation and Language