RoboMIND humanoid dataset surpasses 20 million downloads
The Beijing Innovation Center of Humanoid Robotics says RoboMIND has exceeded 20 million cumulative downloads worldwide, twice the total reported one month earlier. Global Times reported the figure following an announcement by the center. The count is self reported, and the report does not include an independent audit or a figure for unique users.
The RoboMIND humanoid dataset was open sourced in December 2025. It contains more than 300,000 dual arm manipulation trajectories covering over 700 real world tasks, according to the center.
Data collection across 40 robot configurations
RoboMIND is intended to support embodied foundation model training, manipulation policy development and algorithm evaluation. The underlying collection program spans home, retail, industrial and medical settings, although the report does not break down how many trajectories come from each category.
The center operates a training facility of nearly 6,000 square meters in Beijing’s Shijingshan district. It says the site contains more than 30 scenarios, over 150 robot units and 40 robot configurations, allowing the same or similar tasks to be captured across different hardware setups.
Xia Hualin, director of the training base, said it has delivered nearly 30,000 hours of data to external partners. The center reports a data qualification rate above 95 percent, but the article does not define the qualification criteria. More than 70 percent of the facility’s production capacity now serves companies and research institutions working on model training and embodied control systems.
Downloads offer a limited measure of adoption
The 20 million figure shows substantial distribution, but downloads are not equivalent to deployed models, completed training runs or unique developers. Nor can a download count establish whether the trajectories transfer effectively between different humanoid platforms. Task success rates and cross platform evaluations would provide stronger evidence of practical value.
The scale of the wider data shortage remains difficult to pin down. Global Times cited a GF Securities report, referenced by People’s Daily, estimating that embodied models suitable for practical deployment require at least 10 million hours of multimodal interaction data. The same estimate placed current global accumulation below 5 percent of that level.
China is rapidly expanding physical data collection capacity. Citing Xinhua and a report from the China Academy of Information and Communications Technology, Global Times said more than 70 training grounds had entered operation nationwide by the end of June, with another 46 planned or under construction.
One industry estimate cited in the report put China’s stock of compliant real world physical interaction data at about 500,000 hours, compared with tens of millions of hours considered necessary for commercialization. Definitions of compliant or high quality data are not yet standardized, making direct comparisons between facilities and datasets difficult.
Source: globaltimes.cn
