PRAETORIAN: A New Defense Mechanism Against GNN Backdoor Attacks
Researchers have introduced PRAETORIAN, a novel defense strategy designed to protect Graph Neural Networks (GNNs) from backdoor attacks. Unlike previous methods that rely on inspecting specific subgraph patterns or node features, which adaptive attackers can easily circumvent, PRAETORIAN targets the intrinsic requirements of effective backdoors. The system operates on the observation that flipping a victim node's prediction requires substantial influence, forcing attackers to either inject many trigger nodes or use highly influential ones. PRAETORIAN analyzes internal correlations within potential trigger subgraphs to detect abnormally large structures and quantifies external node influence to identify triggers with disproportionate impact. Evaluations show that PRAETORIAN reduces the average attack success rate to 0.55% with only a 0.62% drop in clean accuracy, significantly outperforming state-of-the-art defenses. Furthermore, it forces adversaries into an unfavorable trade-off: achieving high attack success rates results in significant accuracy drops, while preserving accuracy limits attack effectiveness. This research highlights a robust approach to enhancing the security of relational data learning tools against sophisticated adversarial threats.
Wire timeline
PRAETORIAN: A New Defense Mechanism Against GNN Backdoor Attacks
Researchers have introduced PRAETORIAN, a novel defense strategy designed to protect Graph Neural Networks (GNNs) from backdoor attacks. Unlike previous methods that rely on inspecting specific subgraph patterns or node features, which adaptive attackers can easily circumvent, PRAETORIAN targets the intrinsic requirements of effective backdoors. The system operates on the observation that flipping a victim node's prediction requires substantial influence, forcing attackers to either inject many trigger nodes or use highly influential ones. PRAETORIAN analyzes internal correlations within potential trigger subgraphs to detect abnormally large structures and quantifies external node influence to identify triggers with disproportionate impact. Evaluations show that PRAETORIAN reduces the average attack success rate to 0.55% with only a 0.62% drop in clean accuracy, significantly outperforming state-of-the-art defenses. Furthermore, it forces adversaries into an unfavorable trade-off: achieving high attack success rates results in significant accuracy drops, while preserving accuracy limits attack effectiveness. This research highlights a robust approach to enhancing the security of relational data learning tools against sophisticated adversarial threats.
cs.AI updates on arXiv.org