Unitree unveils humanoid robot foundation model, leads open-source benchmarks
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
Unitree Robotics has unveiled its general-purpose humanoid robot foundation model, UnifoLM-ER-1, which achieved top-tier results among open-source models in seven embodied reasoning benchmarks, including spatial understanding and reasoning tasks. The model, built on Qwen3-VL-4B, outperformed other open-source models in benchmarks such as RefSpatial-Bench, Where2Place, and PixMo-Point, with some metrics exceeding closed-source models like Gemini-ER 2. Additionally, Unitree introduced UnifoLM-WLA-1.0, a unified World-Language-Action model covering 64 tasks (10 whole-body operations and 54 tabletop manipulations), integrating embodied reasoning, future dynamic region prediction, and action generation. Unitree has released partial model weights and data, with plans to open post-training code and additional datasets. The company aims to lower development costs for small and medium-sized robotics firms and research institutions, fostering a developer ecosystem. The article notes that competition in humanoid robotics is expected to shift from hardware to foundation models, development interfaces, and application ecosystems, though real-world generalization remains to be validated.
Source report
Recently, Unitree Robotics unveiled its general-purpose humanoid robot. The results indicate that Unitree has entered the top tier of open-source models in terms of embodied spatial perception and reasoning. As model weights and data are gradually opened up, competition in the humanoid robotics industry will extend from hardware performance to foundation models, development interfaces, and application ecosystems.
Leadership in 7 Embodied Reasoning Benchmarks
In 16 multimodal perception and understanding benchmark tests, Unitree's UnifoLM-ER-1 (built on Qwen3-VL-4B) achieved the best results among open-source models in 7 benchmarks, with overall performance comparable to leading closed-source models.
According to the table results, UnifoLM-ER-1's advantages are primarily concentrated in spatial understanding and reasoning. It outperforms the open-source models evaluated across seven benchmarks:
- RefSpatial-Bench
- Where2Place
- PixMo-Point
- BLINK
- EmbSpatial
- RoboSpatial
- VSR
Some metrics also exceed those of closed-source models such as Gemini-ER 2.
These results demonstrate that UnifoLM-ER-1 ranks at the forefront of open-source models in spatial understanding and reasoning. This capability determines whether a robot can identify targets, understand spatial relationships, and provide a basis for subsequent action generation.
The significance of these seven results lies not in whether the robot can stably execute actions—embodied reasoning provides the foundation for robots to judge target locations, plan interaction processes, and generate subsequent actions.
This set of results covers multiple dimensions, including object reference, placement location, visual pointing, and spatial relationships, reflecting the model's relatively comprehensive spatial perception and reasoning capabilities. The open model and evaluation results also provide a basis for future reproduction and comparison.
From this set of benchmark results, it is evident that Unitree's embodied reasoning model has entered the top tier of open-source models.
One Model Covers 64 Tasks: From Tabletop Manipulation to Whole-Body Operation
UnifoLM-WLA-1.0 covers 64 tasks, including 10 whole-body operations and 54 tabletop manipulations.
The focus of these 64 tasks is not merely their quantity, but rather that a single model covers both tabletop and whole-body operations, coordinating arm movements, end-effectors, and lower-limb actions.
This indicates that the same model can cover multiple types of tasks and adapt to various end-effectors included in training and testing. In whole-body tasks such as organizing shoes or working in kitchens or bathrooms, the robot needs to move while performing continuous operations. When making a bed, the dual-arm movements must also coordinate with body posture and lower-limb balance.
WLA in UnifoLM-WLA-1.0 stands for World-Language-Action. The model integrates embodied reasoning, future dynamic region prediction, and action generation into a unified framework, enabling the robot to first understand the current environment and potential changes before generating actions.
Such a unified model is expected to reduce the repetitive work of developing separate strategies for each task and improve the reuse efficiency of the model across similar tasks and supported hardware. Compared to single-task demonstrations, the 64 tasks showcase the ability of one model to cover diverse operations. However, cross-scenario and cross-platform generalization still need to be validated through more real-world environment testing.
Gradually Opening Up Models and Data to Compete for Developer Ecosystems
Currently, Unitree has launched the project homepage for UnifoLM-WLA-1.0 and released partial model weights and data. Post-training code, UnifoLM-WLA-Base weights, whole-body teleoperation datasets, and manipulation datasets remain listed in future opening plans.
After the model is opened, development interfaces and hardware adaptation will influence its scope of application. Embodied models need to integrate with different robot platforms and sensors.
Unitree develops both robot hardware and trains foundation models. Opening up the model can lower trial and secondary development costs for developers, attract more teams to adapt to its model interface, and accelerate iteration through external testing and feedback. These accumulations may expand the influence of its models and interfaces within the industry.
In the short term, open weights and data can reduce the cost for small and medium-sized robotics companies and research institutions to train foundation models from scratch, enabling them to conduct scenario-specific adaptation and validation more quickly. Actual deployment still requires engineering development combined with hardware, data, and control systems.
In the long run, competition for embodied foundation models may intensify further. Some companies will provide foundation models and development platforms adaptable to multiple hardware platforms, while others will focus on robot hardware, industry data, and scenario integration. These two capabilities may also develop synergistically within the same company.
What needs to be observed next is whether the opened models can maintain task success rates and generalization capabilities in more real-world environments, and whether developers can form sustained reproduction, adaptation, and application around the models.
(Source: Red Star Capital Bureau)
Source
东方财富网-A股公司Eastern
Part of this Story
Unitree open-sources humanoid robot foundation model, leading seven benchmarks