LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models
Researchers have introduced LoopVLA, a novel recurrent architecture for Vision-Language-Action (VLA) models designed to enhance robotic manipulation efficiency. Unlike traditional VLA models that rely on deep, fixed-layer representations, LoopVLA iteratively refines multimodal tokens using a shared Transformer block. It jointly learns representation refinement, action prediction, and sufficiency estimation, allowing the system to determine when a representation is adequate for action without unnecessary computation. To address the lack of direct supervision for sufficiency, the model employs a self-supervised distribution alignment objective, linking confidence scores to relative action quality. Experimental results on LIBERO, LIBERO-Plus, and VLA-Arena benchmarks demonstrate significant improvements: LoopVLA reduces model parameters by 45% and increases inference throughput by up to 1.7 times while maintaining or exceeding the task success rates of strong baselines. This approach effectively decouples refinement from absolute layer indices, optimizing both computational resources and control precision for closed-loop spatial adjustments in robotics.
Wire timeline
LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models
Researchers have introduced LoopVLA, a novel recurrent architecture for Vision-Language-Action (VLA) models designed to enhance robotic manipulation efficiency. Unlike traditional VLA models that rely on deep, fixed-layer representations, LoopVLA iteratively refines multimodal tokens using a shared Transformer block. It jointly learns representation refinement, action prediction, and sufficiency estimation, allowing the system to determine when a representation is adequate for action without unnecessary computation. To address the lack of direct supervision for sufficiency, the model employs a self-supervised distribution alignment objective, linking confidence scores to relative action quality. Experimental results on LIBERO, LIBERO-Plus, and VLA-Arena benchmarks demonstrate significant improvements: LoopVLA reduces model parameters by 45% and increases inference throughput by up to 1.7 times while maintaining or exceeding the task success rates of strong baselines. This approach effectively decouples refinement from absolute layer indices, optimizing both computational resources and control precision for closed-loop spatial adjustments in robotics.
cs.AI updates on arXiv.org