Supervised Mixture-of-Experts for Surgical Grasping and Retraction
Researchers have introduced a supervised Mixture-of-Experts (MoE) architecture designed to enhance robotic surgical manipulation, specifically addressing challenges like data scarcity and safety requirements. This framework can be integrated with existing autonomous policies, such as the Action Chunking Transformer (ACT). The study demonstrates that this lightweight approach enables robots to learn complex, long-horizon tasks like bowel grasping and retraction using fewer than 150 demonstrations and only stereo endoscopic images. Unlike generalist Vision Language Action models, which failed entirely, the MoE-enhanced ACT achieved high success rates in standard conditions and showed superior robustness in out-of-distribution scenarios, including novel grasp locations and partial occlusions. Notably, the system generalized to unseen viewpoints and transferred zero-shot to ex vivo porcine tissue without additional training. Preliminary qualitative results from in vivo porcine surgery further support its potential for clinical deployment. This advancement offers a promising pathway for safer, more efficient robotic assistance in surgical environments, reducing the reliance on extensive datasets and multi-camera setups typically required in prior approaches.
Wire timeline
Supervised Mixture-of-Experts for Surgical Grasping and Retraction
Researchers have introduced a supervised Mixture-of-Experts (MoE) architecture designed to enhance robotic surgical manipulation, specifically addressing challenges like data scarcity and safety requirements. This framework can be integrated with existing autonomous policies, such as the Action Chunking Transformer (ACT). The study demonstrates that this lightweight approach enables robots to learn complex, long-horizon tasks like bowel grasping and retraction using fewer than 150 demonstrations and only stereo endoscopic images. Unlike generalist Vision Language Action models, which failed entirely, the MoE-enhanced ACT achieved high success rates in standard conditions and showed superior robustness in out-of-distribution scenarios, including novel grasp locations and partial occlusions. Notably, the system generalized to unseen viewpoints and transferred zero-shot to ex vivo porcine tissue without additional training. Preliminary qualitative results from in vivo porcine surgery further support its potential for clinical deployment. This advancement offers a promising pathway for safer, more efficient robotic assistance in surgical environments, reducing the reliance on extensive datasets and multi-camera setups typically required in prior approaches.
cs.AI updates on arXiv.org