VECTOR-Drive: Tightly Coupled Vision-Language and Trajectory Expert Routing for End-to-End Autonomous Driving
Researchers have introduced VECTOR-DRIVE, a novel end-to-end autonomous driving framework designed to resolve the coupling trade-off in vision-language-action (VLA) models. Built on Qwen2.5-VL-3B, this system tightly integrates semantic reasoning and motion planning within a single multimodal Transformer. Unlike previous approaches that either entangle tasks or weaken semantic-motion links, VECTOR-DRIVE utilizes shared self-attention to keep all tokens coupled while routing feed-forward computation based on token semantics. Vision and language tokens are processed by a Vision-Language Expert to preserve semantic priors, whereas trajectory-related tokens are handled by a Trajectory Expert. A flow-matching planner further refines action tokens into precise waypoints and speed profiles. Evaluations on the Bench2Drive benchmark demonstrate superior performance, with VECTOR-DRIVE achieving a Driving Score of 88.91, outperforming existing end-to-end and VLA-based baselines. The study validates the effectiveness of its architecture through qualitative results and ablation studies, highlighting benefits in shared attention, expert routing, and progressive training for enhanced autonomous driving capabilities.
Wire timeline
VECTOR-Drive: Tightly Coupled Vision-Language and Trajectory Expert Routing for End-to-End Autonomous Driving
Researchers have introduced VECTOR-DRIVE, a novel end-to-end autonomous driving framework designed to resolve the coupling trade-off in vision-language-action (VLA) models. Built on Qwen2.5-VL-3B, this system tightly integrates semantic reasoning and motion planning within a single multimodal Transformer. Unlike previous approaches that either entangle tasks or weaken semantic-motion links, VECTOR-DRIVE utilizes shared self-attention to keep all tokens coupled while routing feed-forward computation based on token semantics. Vision and language tokens are processed by a Vision-Language Expert to preserve semantic priors, whereas trajectory-related tokens are handled by a Trajectory Expert. A flow-matching planner further refines action tokens into precise waypoints and speed profiles. Evaluations on the Bench2Drive benchmark demonstrate superior performance, with VECTOR-DRIVE achieving a Driving Score of 88.91, outperforming existing end-to-end and VLA-based baselines. The study validates the effectiveness of its architecture through qualitative results and ablation studies, highlighting benefits in shared attention, expert routing, and progressive training for enhanced autonomous driving capabilities.
cs.AI updates on arXiv.org