Reflective Test-Time Planning for Embodied LLMs
Researchers have introduced Reflective Test-Time Planning, a novel framework designed to enhance Embodied Large Language Models (LLMs) in robotics. Current embodied LLMs often fail to learn from errors, resulting in repeated mistakes during task execution. This new approach integrates two reflection modes: 'reflection-in-action,' which uses test-time scaling to evaluate candidate actions before execution, and 'reflection-on-action,' which updates the model and policy based on post-execution feedback. Additionally, retrospective reflection allows for long-horizon credit assignment by re-evaluating past decisions. Experiments conducted on newly designed benchmarks, including Long-Horizon Household and MuJoCo Cupboard Fitting tasks, demonstrate significant performance improvements over baseline models. The system also exhibits zero-shot generalization in photorealistic HM3D environments and successful real-world deployment on a Franka Panda robot arm. Ablation studies confirm the mutual dependence of the reflection modes and highlight the efficiency of retrospective reflection compared to step-wise feedback. This advancement enables robots to accumulate experience and correct behaviors through reflective learning rather than independent trials.
Wire timeline
Reflective Test-Time Planning for Embodied LLMs
Researchers have introduced Reflective Test-Time Planning, a novel framework designed to enhance Embodied Large Language Models (LLMs) in robotics. Current embodied LLMs often fail to learn from errors, resulting in repeated mistakes during task execution. This new approach integrates two reflection modes: 'reflection-in-action,' which uses test-time scaling to evaluate candidate actions before execution, and 'reflection-on-action,' which updates the model and policy based on post-execution feedback. Additionally, retrospective reflection allows for long-horizon credit assignment by re-evaluating past decisions. Experiments conducted on newly designed benchmarks, including Long-Horizon Household and MuJoCo Cupboard Fitting tasks, demonstrate significant performance improvements over baseline models. The system also exhibits zero-shot generalization in photorealistic HM3D environments and successful real-world deployment on a Franka Panda robot arm. Ablation studies confirm the mutual dependence of the reflection modes and highlight the efficiency of retrospective reflection compared to step-wise feedback. This advancement enables robots to accumulate experience and correct behaviors through reflective learning rather than independent trials.
cs.AI updates on arXiv.org