Hierarchical Reasoning for Continuous Motion in Closed-Loop Traffic Simulation
Researchers have proposed a novel hierarchical architecture for closed-loop traffic simulation, aiming to overcome the limitations of recent self-play reinforcement learning approaches. While self-play methods offer scalability, they often fail to replicate the socially aware behaviors characteristic of human drivers. The new framework combines high-level multi-agent interaction reasoning with low-level continuous trajectory realization. Specifically, it utilizes a Stackelberg-style Multi-Agent Reinforcement Learning (MARL) module to generate interaction-aware intention commands. These commands guide a low-level continuous motion module, which translates strategic intent into physically consistent and scene-responsive control sequences. To address distribution shift issues during closed-loop deployment, the team introduced a hybrid co-training scheme that integrates MARL with auxiliary recovery supervision. Experimental results conducted on a SUMO-based urban network indicate that this approach achieves superior control smoothness and safety compared to both self-play and passive imitation baselines, while maintaining competitive traffic efficiency. This advancement represents a significant step toward creating more realistic and scalable agents for autonomous driving simulations.
Wire timeline
Hierarchical Reasoning for Continuous Motion in Closed-Loop Traffic Simulation
Researchers have proposed a novel hierarchical architecture for closed-loop traffic simulation, aiming to overcome the limitations of recent self-play reinforcement learning approaches. While self-play methods offer scalability, they often fail to replicate the socially aware behaviors characteristic of human drivers. The new framework combines high-level multi-agent interaction reasoning with low-level continuous trajectory realization. Specifically, it utilizes a Stackelberg-style Multi-Agent Reinforcement Learning (MARL) module to generate interaction-aware intention commands. These commands guide a low-level continuous motion module, which translates strategic intent into physically consistent and scene-responsive control sequences. To address distribution shift issues during closed-loop deployment, the team introduced a hybrid co-training scheme that integrates MARL with auxiliary recovery supervision. Experimental results conducted on a SUMO-based urban network indicate that this approach achieves superior control smoothness and safety compared to both self-play and passive imitation baselines, while maintaining competitive traffic efficiency. This advancement represents a significant step toward creating more realistic and scalable agents for autonomous driving simulations.
cs.AI updates on arXiv.org