RePO-VLA: Recovery-Driven Policy Optimization for Vision-Language-Action Models
Researchers have introduced RePO-VLA, a novel framework designed to enhance the robustness of Vision-Language-Action (VLA) models in complex, long-horizon manipulation tasks. Traditional VLA models often fail due to execution drift and the discarding of failed training data. RePO-VLA addresses this by utilizing success, recovery, and failure trajectories distinctively. It employs Recovery-Aware Initialization to reset history during corrective actions and learns a Progress-Aware Semantic Value Function to align trajectory features with instructions. This approach salvages useful data from failures and teaches the model to distinguish between nominal, failed, and corrective actions. Additionally, a Value-Conditioned Refinement process biases the policy toward high-progress actions. The team also launched FRBench, a benchmark for evaluating recovery capabilities. Experimental results demonstrate significant improvements, with adversarial success rates rising from 20% to 75% in simulations and up to 80% in real-world bimanual tasks, marking a substantial advancement in robotic manipulation reliability without requiring online failure detectors.
Wire timeline
RePO-VLA: Recovery-Driven Policy Optimization for Vision-Language-Action Models
Researchers have introduced RePO-VLA, a novel framework designed to enhance the robustness of Vision-Language-Action (VLA) models in complex, long-horizon manipulation tasks. Traditional VLA models often fail due to execution drift and the discarding of failed training data. RePO-VLA addresses this by utilizing success, recovery, and failure trajectories distinctively. It employs Recovery-Aware Initialization to reset history during corrective actions and learns a Progress-Aware Semantic Value Function to align trajectory features with instructions. This approach salvages useful data from failures and teaches the model to distinguish between nominal, failed, and corrective actions. Additionally, a Value-Conditioned Refinement process biases the policy toward high-progress actions. The team also launched FRBench, a benchmark for evaluating recovery capabilities. Experimental results demonstrate significant improvements, with adversarial success rates rising from 20% to 75% in simulations and up to 80% in real-world bimanual tasks, marking a substantial advancement in robotic manipulation reliability without requiring online failure detectors.
cs.AI updates on arXiv.org