Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs
Researchers have introduced a novel framework called Retrieve-then-Steer to enhance the reliability of Vision-Language-Action (VLA) models in robotic manipulation. While VLA models show promise, their performance often degrades in specific local deployment conditions. This study addresses persistent-deployment settings where robots operate repeatedly in similar environments. The proposed method utilizes an online success-memory guided test-time adaptation mechanism. During operation, the robot stores successful observation-action segments in long-term memory. At inference, it retrieves relevant action chunks, filters them for consistency, and aggregates them into an elite action prior. This prior is injected into the flow-matching action sampler using confidence-adaptive guidance, allowing the frozen VLA model to leverage past successes without requiring parameter updates. Experiments in both simulation and real-world scenarios demonstrate that this lightweight, non-parametric approach significantly improves task success rates and closed-loop stability, particularly for long-horizon and multi-stage tasks. The research highlights a shift from treating test episodes as independent zero-shot trials to leveraging environment-verified evidence for continuous improvement.
Wire timeline
Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs
Researchers have introduced a novel framework called Retrieve-then-Steer to enhance the reliability of Vision-Language-Action (VLA) models in robotic manipulation. While VLA models show promise, their performance often degrades in specific local deployment conditions. This study addresses persistent-deployment settings where robots operate repeatedly in similar environments. The proposed method utilizes an online success-memory guided test-time adaptation mechanism. During operation, the robot stores successful observation-action segments in long-term memory. At inference, it retrieves relevant action chunks, filters them for consistency, and aggregates them into an elite action prior. This prior is injected into the flow-matching action sampler using confidence-adaptive guidance, allowing the frozen VLA model to leverage past successes without requiring parameter updates. Experiments in both simulation and real-world scenarios demonstrate that this lightweight, non-parametric approach significantly improves task success rates and closed-loop stability, particularly for long-horizon and multi-stage tasks. The research highlights a shift from treating test episodes as independent zero-shot trials to leveraging environment-verified evidence for continuous improvement.
cs.AI updates on arXiv.org