VLADriver-RAG: Retrieval-Augmented Vision-Language-Action Models for Autonomous Driving
Researchers have introduced VLADriver-RAG, a novel framework designed to enhance end-to-end autonomous driving by integrating Retrieval-Augmented Generation (RAG) with Vision-Language-Action (VLA) models. While VLA models show promise, they often struggle with generalization in rare, long-tail scenarios due to reliance on implicit knowledge. Standard visual retrieval methods also face issues with high latency and semantic ambiguity. To overcome these limitations, VLADriver-RAG grounds planning in explicit, structure-aware historical knowledge. It employs a Visual-to-Scenario mechanism to convert sensory inputs into spatiotemporal semantic graphs, effectively filtering visual noise. Additionally, a Scenario-Aligned Embedding Model uses Graph-DTW metric alignment to prioritize topological consistency over superficial visual similarity, ensuring higher retrieval relevance. These retrieved priors are fused into a query-based VLA backbone to generate precise trajectories. Extensive testing on the Bench2Drive benchmark demonstrates that this approach achieves a new state-of-the-art Driving Score of 89.12, marking a significant advancement in autonomous driving technology and addressing critical challenges in scenario generalization and planning accuracy.
Wire timeline
VLADriver-RAG: Retrieval-Augmented Vision-Language-Action Models for Autonomous Driving
Researchers have introduced VLADriver-RAG, a novel framework designed to enhance end-to-end autonomous driving by integrating Retrieval-Augmented Generation (RAG) with Vision-Language-Action (VLA) models. While VLA models show promise, they often struggle with generalization in rare, long-tail scenarios due to reliance on implicit knowledge. Standard visual retrieval methods also face issues with high latency and semantic ambiguity. To overcome these limitations, VLADriver-RAG grounds planning in explicit, structure-aware historical knowledge. It employs a Visual-to-Scenario mechanism to convert sensory inputs into spatiotemporal semantic graphs, effectively filtering visual noise. Additionally, a Scenario-Aligned Embedding Model uses Graph-DTW metric alignment to prioritize topological consistency over superficial visual similarity, ensuring higher retrieval relevance. These retrieved priors are fused into a query-based VLA backbone to generate precise trajectories. Extensive testing on the Bench2Drive benchmark demonstrates that this approach achieves a new state-of-the-art Driving Score of 89.12, marking a significant advancement in autonomous driving technology and addressing critical challenges in scenario generalization and planning accuracy.
cs.AI updates on arXiv.org