BehaviorGuard: Online Backdoor Defense for Deep Reinforcement Learning
Researchers have introduced BehaviorGuard, a novel online framework designed to defend Deep Reinforcement Learning (DRL) systems against backdoor attacks. Unlike traditional defenses that rely on reverse-engineering triggers or costly model fine-tuning, BehaviorGuard focuses on trigger-agnostic output behaviors. The study reveals that backdoored policies cause consistent shifts in action distributions, leaving detectable traces in high-quantile regions and distribution tails even without active triggers. By utilizing a new metric to capture this behavioral drift, the framework identifies and suppresses malicious actions at runtime. This approach is significant as it represents the first online defense capable of countering attacks in both single-agent and multi-agent DRL environments. Evaluations across diverse benchmarks demonstrate that BehaviorGuard consistently outperforms existing methods in terms of both efficacy and efficiency, offering a more robust and practical solution for securing AI agents against sophisticated security threats.
Wire timeline
BehaviorGuard: Online Backdoor Defense for Deep Reinforcement Learning
Researchers have introduced BehaviorGuard, a novel online framework designed to defend Deep Reinforcement Learning (DRL) systems against backdoor attacks. Unlike traditional defenses that rely on reverse-engineering triggers or costly model fine-tuning, BehaviorGuard focuses on trigger-agnostic output behaviors. The study reveals that backdoored policies cause consistent shifts in action distributions, leaving detectable traces in high-quantile regions and distribution tails even without active triggers. By utilizing a new metric to capture this behavioral drift, the framework identifies and suppresses malicious actions at runtime. This approach is significant as it represents the first online defense capable of countering attacks in both single-agent and multi-agent DRL environments. Evaluations across diverse benchmarks demonstrate that BehaviorGuard consistently outperforms existing methods in terms of both efficacy and efficiency, offering a more robust and practical solution for securing AI agents against sophisticated security threats.
cs.AI updates on arXiv.org