AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models
Researchers have introduced AsyncVLA, a novel framework designed to enhance Vision-Language-Action (VLA) models for generalist robotics. Traditional VLA models rely on synchronous flow matching with rigid time schedules, which often leads to instability and cascading errors in long-horizon tasks. AsyncVLA addresses these limitations by implementing asynchronous flow matching, allowing for non-uniform time schedules and action context awareness. A key innovation is the inclusion of a confidence rater that evaluates initially generated actions, enabling the model to selectively refine inaccurate tokens before execution. The framework also features a unified training procedure for both synchronous and asynchronous modes, improving KV-cache utilization. Extensive experiments on robotic manipulation benchmarks demonstrate that AsyncVLA is data-efficient and possesses strong self-correction capabilities. The model outperforms existing methods in both simulation and real-world evaluations, marking a significant advancement in robust robotic control. The source code has been made publicly available to support further research and development in this field.
Wire timeline
AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models
Researchers have introduced AsyncVLA, a novel framework designed to enhance Vision-Language-Action (VLA) models for generalist robotics. Traditional VLA models rely on synchronous flow matching with rigid time schedules, which often leads to instability and cascading errors in long-horizon tasks. AsyncVLA addresses these limitations by implementing asynchronous flow matching, allowing for non-uniform time schedules and action context awareness. A key innovation is the inclusion of a confidence rater that evaluates initially generated actions, enabling the model to selectively refine inaccurate tokens before execution. The framework also features a unified training procedure for both synchronous and asynchronous modes, improving KV-cache utilization. Extensive experiments on robotic manipulation benchmarks demonstrate that AsyncVLA is data-efficient and possesses strong self-correction capabilities. The model outperforms existing methods in both simulation and real-world evaluations, marking a significant advancement in robust robotic control. The source code has been made publicly available to support further research and development in this field.
cs.AI updates on arXiv.org