Empirical Evaluation of Domain-Adapted Language Models for STRIDE Threat Modelling in 5G Security
A new study published on arXiv evaluates the effectiveness of Large Language Models (LLMs) and Small Language Models (SLMs) in structured threat modelling, specifically within the context of 5G security. The research systematically compares domain-adapted models against general-purpose counterparts using the STRIDE methodology across 52 different configurations involving eight distinct language models. Key variables analyzed include domain adaptation, model scale, decoding strategies, and prompting techniques. The findings reveal that domain-adapted models do not consistently outperform general-purpose models, and while larger models generally show higher performance, the gains are insufficient for reliable threat modelling. Decoding strategies significantly impact output validity. The study highlights fundamental limitations in current LLM capabilities for this task, suggesting that simple scaling or additional training data is inadequate. Instead, the authors advocate for incorporating task-specific reasoning and stronger grounding in security concepts. The paper also provides insights into invalid outputs and offers tailored prompting suggestions for STRIDE threat classification, contributing to the ongoing discourse on AI applications in cybersecurity.
Wire timeline
Empirical Evaluation of Domain-Adapted Language Models for STRIDE Threat Modelling in 5G Security
A new study published on arXiv evaluates the effectiveness of Large Language Models (LLMs) and Small Language Models (SLMs) in structured threat modelling, specifically within the context of 5G security. The research systematically compares domain-adapted models against general-purpose counterparts using the STRIDE methodology across 52 different configurations involving eight distinct language models. Key variables analyzed include domain adaptation, model scale, decoding strategies, and prompting techniques. The findings reveal that domain-adapted models do not consistently outperform general-purpose models, and while larger models generally show higher performance, the gains are insufficient for reliable threat modelling. Decoding strategies significantly impact output validity. The study highlights fundamental limitations in current LLM capabilities for this task, suggesting that simple scaling or additional training data is inadequate. Instead, the authors advocate for incorporating task-specific reasoning and stronger grounding in security concepts. The paper also provides insights into invalid outputs and offers tailored prompting suggestions for STRIDE threat classification, contributing to the ongoing discourse on AI applications in cybersecurity.
cs.AI updates on arXiv.org