AffordSim: A Scalable Data Generator and Benchmark for Affordance-Aware Robotic Manipulation
Researchers have introduced AffordSim, a novel scalable data generator and benchmark designed to enhance affordance-aware robotic manipulation. Current simulation methods often rely on generic grasp estimators that lack task semantics or require labor-intensive manual annotations for each new object. AffordSim addresses these limitations by integrating open-vocabulary 3D affordance prediction into simulation-based trajectory generation. The system synthesizes task-relevant scenes from natural-language descriptions, grounds affordance queries on object surfaces, and selects executable grasps through motion planning. It also incorporates randomization of pose, texture, and lighting to facilitate sim-to-real transfer. Evaluated as a benchmark with 50 tasks, five robot embodiments, and over 500 objects, AffordSim achieves 93% of the success rate of manual annotations on critical tasks. Furthermore, vision-language-action policies trained on AffordSim data demonstrated zero-shot transfer capabilities to a real-world Franka FR3 robot, achieving a 24% average success rate. This development represents a significant step forward in automating high-quality training data generation for complex robotic manipulation skills.
Wire timeline
AffordSim: A Scalable Data Generator and Benchmark for Affordance-Aware Robotic Manipulation
Researchers have introduced AffordSim, a novel scalable data generator and benchmark designed to enhance affordance-aware robotic manipulation. Current simulation methods often rely on generic grasp estimators that lack task semantics or require labor-intensive manual annotations for each new object. AffordSim addresses these limitations by integrating open-vocabulary 3D affordance prediction into simulation-based trajectory generation. The system synthesizes task-relevant scenes from natural-language descriptions, grounds affordance queries on object surfaces, and selects executable grasps through motion planning. It also incorporates randomization of pose, texture, and lighting to facilitate sim-to-real transfer. Evaluated as a benchmark with 50 tasks, five robot embodiments, and over 500 objects, AffordSim achieves 93% of the success rate of manual annotations on critical tasks. Furthermore, vision-language-action policies trained on AffordSim data demonstrated zero-shot transfer capabilities to a real-world Franka FR3 robot, achieving a 24% average success rate. This development represents a significant step forward in automating high-quality training data generation for complex robotic manipulation skills.
cs.AI updates on arXiv.org