CoWorld-VLA: Multi-Expert World Model for Autonomous Driving
Researchers have introduced CoWorld-VLA, a novel Vision-Language-Action (VLA) framework designed to enhance end-to-end autonomous driving systems. Addressing the limitations of existing reasoning mechanisms, such as textual Chain-of-Thought and latent world reasoning, this model employs a multi-expert world reasoning approach. It extracts complementary world information through multi-source supervision, encoding it into four specific expert tokens: semantic interaction, geometric structure, dynamic evolution, and ego trajectory. These tokens serve as explicit conditions to guide action planning, modeling interaction intent, spatial structure, future dynamics, and behavioral goals respectively. The system utilizes a diffusion-based hierarchical multi-expert fusion planner coupled with scene context to generate continuous ego trajectories. Experimental results on the NAVSIM v1 benchmark demonstrate that CoWorld-VLA achieves competitive performance in both future scene generation and planning, showing significant improvements in collision avoidance and trajectory accuracy. Ablation studies confirm the effectiveness and complementarity of the expert tokens. The source code for this framework has been made publicly available to support further research and development in autonomous driving technologies.
Wire timeline
CoWorld-VLA: Multi-Expert World Model for Autonomous Driving
Researchers have introduced CoWorld-VLA, a novel Vision-Language-Action (VLA) framework designed to enhance end-to-end autonomous driving systems. Addressing the limitations of existing reasoning mechanisms, such as textual Chain-of-Thought and latent world reasoning, this model employs a multi-expert world reasoning approach. It extracts complementary world information through multi-source supervision, encoding it into four specific expert tokens: semantic interaction, geometric structure, dynamic evolution, and ego trajectory. These tokens serve as explicit conditions to guide action planning, modeling interaction intent, spatial structure, future dynamics, and behavioral goals respectively. The system utilizes a diffusion-based hierarchical multi-expert fusion planner coupled with scene context to generate continuous ego trajectories. Experimental results on the NAVSIM v1 benchmark demonstrate that CoWorld-VLA achieves competitive performance in both future scene generation and planning, showing significant improvements in collision avoidance and trajectory accuracy. Ablation studies confirm the effectiveness and complementarity of the expert tokens. The source code for this framework has been made publicly available to support further research and development in autonomous driving technologies.
cs.AI updates on arXiv.org