EvoDriveVLA: Evolving Driving VLA Models via Collaborative Perception-Planning Distillation
Researchers have introduced EvoDriveVLA, a novel framework designed to enhance Vision-Language-Action (VLA) models for autonomous driving. Current VLA models often face challenges such as degraded perception after unfreezing visual encoders and instability in long-term planning. To address these issues, EvoDriveVLA employs a collaborative perception-planning distillation approach. This method integrates self-anchored perceptual constraints, using a self-anchor teacher to regularize student representations through trajectory-guided key-region awareness. Additionally, it utilizes future-informed trajectory distillation, where a future-aware oracle teacher synthesizes reasoning trajectories via coarse-to-fine refinement and Monte Carlo dropout sampling. This allows the student model to internalize future-aware insights effectively. The framework has demonstrated state-of-the-art performance in nuScenes open-loop evaluations and significantly improved results in NAVSIM closed-loop assessments. The study highlights advancements in stabilizing visual encoding and improving trajectory prediction for safer autonomous navigation. The associated code has been made publicly available to support further research and development in this field.
Wire timeline
EvoDriveVLA: Evolving Driving VLA Models via Collaborative Perception-Planning Distillation
Researchers have introduced EvoDriveVLA, a novel framework designed to enhance Vision-Language-Action (VLA) models for autonomous driving. Current VLA models often face challenges such as degraded perception after unfreezing visual encoders and instability in long-term planning. To address these issues, EvoDriveVLA employs a collaborative perception-planning distillation approach. This method integrates self-anchored perceptual constraints, using a self-anchor teacher to regularize student representations through trajectory-guided key-region awareness. Additionally, it utilizes future-informed trajectory distillation, where a future-aware oracle teacher synthesizes reasoning trajectories via coarse-to-fine refinement and Monte Carlo dropout sampling. This allows the student model to internalize future-aware insights effectively. The framework has demonstrated state-of-the-art performance in nuScenes open-loop evaluations and significantly improved results in NAVSIM closed-loop assessments. The study highlights advancements in stabilizing visual encoding and improving trajectory prediction for safer autonomous navigation. The associated code has been made publicly available to support further research and development in this field.
cs.AI updates on arXiv.org