ThreatCore: A Benchmark for Explicit and Implicit Threat Detection
Researchers have introduced ThreatCore, a new publicly available benchmark dataset designed to address the lack of consistent definitions and standardized benchmarks in Natural Language Processing threat detection. Often conflated with toxicity or hate speech, threat detection requires finer granularity. ThreatCore distinguishes between explicit threats, implicit threats, and non-threats by aggregating multiple public resources and systematically re-annotating them under a unified operational definition. This process revealed significant inconsistencies in existing labels. To enhance coverage of underrepresented cases, particularly implicit threats, the dataset was augmented with synthetic examples that were manually validated using the same rigorous annotation protocol. Evaluations of Perspective API, zero-shot classifiers, and recent language models on ThreatCore demonstrate that implicit threats remain substantially more difficult to detect than explicit ones. The study also finds that incorporating Semantic Role Labeling as an intermediate representation improves performance by clarifying the structure of harmful intent. Overall, ThreatCore provides a consistent framework for studying fine-grained threat detection and highlights ongoing challenges in identifying indirect expressions of harmful intent in AI systems.
Wire timeline
ThreatCore: A Benchmark for Explicit and Implicit Threat Detection
Researchers have introduced ThreatCore, a new publicly available benchmark dataset designed to address the lack of consistent definitions and standardized benchmarks in Natural Language Processing threat detection. Often conflated with toxicity or hate speech, threat detection requires finer granularity. ThreatCore distinguishes between explicit threats, implicit threats, and non-threats by aggregating multiple public resources and systematically re-annotating them under a unified operational definition. This process revealed significant inconsistencies in existing labels. To enhance coverage of underrepresented cases, particularly implicit threats, the dataset was augmented with synthetic examples that were manually validated using the same rigorous annotation protocol. Evaluations of Perspective API, zero-shot classifiers, and recent language models on ThreatCore demonstrate that implicit threats remain substantially more difficult to detect than explicit ones. The study also finds that incorporating Semantic Role Labeling as an intermediate representation improves performance by clarifying the structure of harmful intent. Overall, ThreatCore provides a consistent framework for studying fine-grained threat detection and highlights ongoing challenges in identifying indirect expressions of harmful intent in AI systems.
cs.AI updates on arXiv.org