Humanoid-LLA: Large Language Action Model for Free-form Humanoid Robot Control
Researchers have introduced Humanoid-LLA, a novel Large Language Action model designed to enable humanoid robots to execute complex, free-form natural language commands. Addressing critical limitations in existing methods, such as data scarcity and physical instability, this approach translates unconstrained language directly into executable whole-body motions. The model employs a unified human-humanoid motion vocabulary to bridge high-level semantics with physically grounded control. Additionally, it utilizes a two-stage fine-tuning framework combining supervised motion Chain-of-Thought learning with reinforcement learning refined by physical feedback. Extensive evaluations in both simulation and real-world cross-embodiment experiments demonstrate that Humanoid-LLA achieves superior generalization to new commands and diverse motion generation while maintaining high physical fidelity. This advancement represents a significant step toward seamless human-robot interaction and general-purpose embodied AI, overcoming previous trade-offs between instruction complexity and motion plausibility.
Wire timeline
Humanoid-LLA: Large Language Action Model for Free-form Humanoid Robot Control
Researchers have introduced Humanoid-LLA, a novel Large Language Action model designed to enable humanoid robots to execute complex, free-form natural language commands. Addressing critical limitations in existing methods, such as data scarcity and physical instability, this approach translates unconstrained language directly into executable whole-body motions. The model employs a unified human-humanoid motion vocabulary to bridge high-level semantics with physically grounded control. Additionally, it utilizes a two-stage fine-tuning framework combining supervised motion Chain-of-Thought learning with reinforcement learning refined by physical feedback. Extensive evaluations in both simulation and real-world cross-embodiment experiments demonstrate that Humanoid-LLA achieves superior generalization to new commands and diverse motion generation while maintaining high physical fidelity. This advancement represents a significant step toward seamless human-robot interaction and general-purpose embodied AI, overcoming previous trade-offs between instruction complexity and motion plausibility.
cs.AI updates on arXiv.org