SCORP: Scene-Consistent Multi-agent Diffusion Planning for Cooperative Driving
Researchers have introduced SCORP, a novel multi-agent diffusion planner designed to enhance safety and efficiency in cooperative driving systems. Addressing the limitations of existing diffusion-based methods, which often struggle with scene consistency and closed-loop alignment, SCORP employs a scene-conditioned denoising architecture. This pre-training phase utilizes inter-agent self-attention and a dual-path conditioning mechanism, combining cross-attention for direct scene information injection with AdaLN-Zero for stable conditional modulation. For post-training, the team formulated a two-layer Markov decision process that integrates the reverse denoising chain with policy-environment interactions. To ensure stability during online reinforcement learning, they developed variance-gated group-relative policy optimization (VG-GRPO) alongside dense planning rewards. Extensive experiments on the WOMD benchmark demonstrate that SCORP significantly outperforms strong open-source baselines, achieving improvements of 10.47%-28.26% in core safety metrics and 1.70%-7.22% in efficiency metrics. These results highlight the model's ability to deliver consistent gains in driving safety and traffic efficiency, marking a substantial advance in stable, closed-loop cooperative autonomous driving technologies.
Wire timeline
SCORP: Scene-Consistent Multi-agent Diffusion Planning for Cooperative Driving
Researchers have introduced SCORP, a novel multi-agent diffusion planner designed to enhance safety and efficiency in cooperative driving systems. Addressing the limitations of existing diffusion-based methods, which often struggle with scene consistency and closed-loop alignment, SCORP employs a scene-conditioned denoising architecture. This pre-training phase utilizes inter-agent self-attention and a dual-path conditioning mechanism, combining cross-attention for direct scene information injection with AdaLN-Zero for stable conditional modulation. For post-training, the team formulated a two-layer Markov decision process that integrates the reverse denoising chain with policy-environment interactions. To ensure stability during online reinforcement learning, they developed variance-gated group-relative policy optimization (VG-GRPO) alongside dense planning rewards. Extensive experiments on the WOMD benchmark demonstrate that SCORP significantly outperforms strong open-source baselines, achieving improvements of 10.47%-28.26% in core safety metrics and 1.70%-7.22% in efficiency metrics. These results highlight the model's ability to deliver consistent gains in driving safety and traffic efficiency, marking a substantial advance in stable, closed-loop cooperative autonomous driving technologies.
cs.AI updates on arXiv.org