Benchmarking Safety Risks of Knowledge-Intensive Reasoning under Malicious Knowledge Editing
Researchers have introduced EditRisk-Bench, a new benchmark designed to systematically evaluate the safety risks associated with knowledge-intensive reasoning in Large Language Models (LLMs) when subjected to malicious knowledge editing. While LLMs increasingly rely on knowledge editing for flexibility, this feature exposes them to adversarial attacks where injected misinformation or bias can corrupt downstream reasoning and lead to harmful outcomes. Unlike previous benchmarks that focused primarily on editing efficacy, EditRisk-Bench assesses how injected knowledge impacts reasoning behavior and reliability across diverse malicious scenarios, including safety violations and bias. Extensive experiments on both open-source and closed-source LLMs demonstrate that malicious editing can reliably induce incorrect or unsafe reasoning while preserving general capabilities, making these risks difficult to detect. The study identifies key influencing factors such as edit scale, knowledge characteristics, and reasoning complexity. This new framework provides an extensible testbed for understanding and mitigating security vulnerabilities in LLM knowledge editing, addressing a critical gap in current AI safety evaluation methodologies.
Wire timeline
Benchmarking Safety Risks of Knowledge-Intensive Reasoning under Malicious Knowledge Editing
Researchers have introduced EditRisk-Bench, a new benchmark designed to systematically evaluate the safety risks associated with knowledge-intensive reasoning in Large Language Models (LLMs) when subjected to malicious knowledge editing. While LLMs increasingly rely on knowledge editing for flexibility, this feature exposes them to adversarial attacks where injected misinformation or bias can corrupt downstream reasoning and lead to harmful outcomes. Unlike previous benchmarks that focused primarily on editing efficacy, EditRisk-Bench assesses how injected knowledge impacts reasoning behavior and reliability across diverse malicious scenarios, including safety violations and bias. Extensive experiments on both open-source and closed-source LLMs demonstrate that malicious editing can reliably induce incorrect or unsafe reasoning while preserving general capabilities, making these risks difficult to detect. The study identifies key influencing factors such as edit scale, knowledge characteristics, and reasoning complexity. This new framework provides an extensible testbed for understanding and mitigating security vulnerabilities in LLM knowledge editing, addressing a critical gap in current AI safety evaluation methodologies.
cs.AI updates on arXiv.org