Ten questions to the humanoid industry
Ten open questions. Ten hypotheses. And, for each one, the evidence sitting behind it.
The demos look finished. Two arms, a box, a clean grab, and a caption that says autonomous. Watch enough of them and you start to believe the hard part is over.
It isn’t. Strip the shipment figures down to what’s actually doing paid work, price what it costs to teach a single task, and ask a supplier for hours in production rather than units sold – and a very different picture shows up. That’s what these ten questions are for. A disclosure while we’re here: this analysis was first presented by Lars-Fredrik Forberg, who is also CEO of Kinetic Blocks, a company in the training-data market that questions four and five touch directly. Weigh what follows with that in mind.
How many humanoid robots are actually doing paid industrial work today?
Four independent sources now let you build the answer from the bottom up. It lands in the hundreds – not the tens of thousands the shipment numbers imply.
Accenture put 19,000 machines out in the first half of the year and said most of it is still proof of concept, internal testing or education. Up to 70 percent of Chinese output in that window went to training centres – roughly 13,000 machines across 53 facilities, with 34 more under construction. IDC reached the same place from another direction: more than 85 percent of the 2025 base sat in performances, education, data collection and guided tours.
You can count the named industrial deployments on your hands. Agility at GXO. Apptronik at Mercedes and Jabil. Figure at BMW. Humanoid at Schaeffler from late 2026. Wandercraft and Renault, with 350 units planned before the end of 2027. Unitree’s founder puts a deployed machine at 30 to 50 percent of a human worker’s task efficiency. And of the nine named Chinese deployments, nine run on a fixed scene, a single task, usually overnight. The world gets arranged for the machine, not the other way around.
Sources: Accenture, Humanoid Robots Summit, Stuttgart, 9 September 2026; The Economist, 23 August 2026, citing Interact Analysis; IDC 2025 humanoid shipments; Core Matter, 24 August 2026.
If the models are this good, what still doesn’t work?
Manipulation looks close to solved in the lab. On real hardware the best of ten policies reaches 12.8 percent, and European safety testing puts almost every collision result in the yellow-to-red band.
Real-world task success rate by policy and robot embodiment — RoboDojo leaderboard, percent
| Policy | ARX X5 | Piper | Piper X | All 18 tasks |
|---|---|---|---|---|
| π0.5 | 13.3 | 21.7 | 3.3 | 12.8 |
| InternVLA-A1 | 3.3 | 18.3 | 0 | 7.2 |
| GalaxeaVLA | 0 | 13.3 | 0 | 4.4 |
| Xiaomi-Robotics-0 | 8.3 | 3.3 | 0 | 3.9 |
| X-VLA | 1.7 | 8.3 | 0 | 3.3 |
| GR00T-N1.7 | 1.7 | 3.3 | 0 | 1.7 |
| π0 | 0 | 5.0 | 0 | 1.7 |
| StarVLA-α | 0 | 5.0 | 0 | 1.7 |
| Spirit v1.5 | 0 | 1.7 | 0 | 0.6 |
| DM0 | 0 | 0 | 0 | 0 |
| Human teleoperation | 100 | 100 | 100 | 100 |
Start with the table. RoboDojo took ten leading policies and put them on real hardware – eighteen real tasks, three different robot arms. The best of them finishes the job 12.8 percent of the time. Look down the Piper X column: nine of the ten complete nothing at all. Not low. Zero. And the bottom row, human operators on the same tasks and the same machines, is a clean 100. The machines can do the work. The models can’t, yet.
Then the benchmark everyone quotes. LIBERO reads as solved at 95 to 99 percent. LIBERO-Plus nudges the camera inside the same simulator and the same models fall below 30 – so this isn’t the sim-to-real gap, it’s models leaning on a view they were handed. Europe has its own version of the finding: Fraunhofer IPA ran Unitree G1, Dobot Atom, Agibot G2 and Neura machines through 66 tests, and the G1 at 50 kilograms isn’t permitted for power-and-force-limited operation under EU norms. The problem isn’t capability. It’s robustness.
Sources: LIBERO-Plus, arXiv 2510.13626; RoboDojo real-world leaderboard, table 2, 3 July 2026; Fraunhofer IPA benchmarking programme, Stuttgart, September 2026; Core Matter, 24 August 2026.
The sim-and-real benchmark behind those success rates
What does it cost to teach a robot one task?
Three separate stages in Stuttgart, on one day, put the price of teaching a single task at four hours, at five to ten hours, and at six months. That spread is the distance between a demonstration and an industrial line.
Retasking a humanoid today takes about three hours against the five minutes the VDMA says industry actually needs. Kitting is still unsolved – dexterous hands broke on metal parts within days. Every speaker described the same shortage: not data in general, but the right data, recorded on operating industrial lines. Rhoda post-trains on one to ten hours of trajectory data per task. That’s the low end of the same range, and it’s the number every buyer should be quoting against.
Sources: Werner Kraus, Fraunhofer IPA; Aya Durbin, Boston Dynamics; Laurent Duthoit, Renault and Camille Croze, Wandercraft; Patrick Schwarzkopf, VDMA – all Humanoid Robots Summit, Stuttgart, September 2026.
Calvin, the new-generation robot, at work on the line
Should you buy your training data, build the channel, or take it for nothing?
This month Figure and Skild both built their own collection channels while LightWheel released 100,000 hours free on Hugging Face. The vendor market is being squeezed from both directions at once.
Buy, build, or take. Real machine data runs 500 to 700 yuan an hour – call it 70 to 98 dollars – against 3 to 40 for egocentric human video. Build it yourself and you’re looking at Figure’s billion-dollar-plus, twelve-month data and compute commitment. Skild spends three dollars on quality control for every one on collection; Chinese training centres get three usable hours out of every eight recorded. And underneath all of it sits general web video – free, effectively unlimited, and, as question five shows, now proven to move a real industrial task. What stays scarce is the demonstration recorded on the customer’s own line.
Sources: Figure, “Introducing Index”, 25 August 2026; Skild AI, “Introducing S1”, 18 August 2026; LightWheel EgoSuite-Open100K via Core Matter; Corey Chan, HSBC, via The Economist, 23 August 2026; Rhoda AI.
Can a robot learn to work by watching video that has nothing to do with robots?
Scaling pre-training on general web video – no actions, almost nothing about robots – took at-speed completion on a real customer task from 4 percent to 85 across seven measured checkpoints.
Task performance against pre-training quality, seven checkpoints across model size and compute, percent
Here’s the part that’s rare in this sector: every point is measured, not modelled. How well a model predicts held-out web video – gauged before it has seen the task at all – ranked the seven checkpoints in the exact order the robot then did. The task is unpacking bearings at a real customer site, over a thousand boxes a day by hand today. General web video carries no actions whatsoever, which puts a free and effectively unlimited layer underneath the entire training-data market. The honest caveat: human and robot limbs move differently, usable footage still has to be labelled by hand, and this is one architecture family on one task.
Sources: Rhoda AI, “Does Scaling Web-Video Pre-training Help Real Robots Do Real Work?”, 10 September 2026; Dyna Robotics, Dyna-2, August 2026. Rhoda states the relationship is a correlation across seven checkpoints.
Rhoda on the task the seven checkpoints were measured on
How should you tell a robot what you want it to do?
On tasks the model has never seen, one video demonstration in context reaches 66 percent where language prompting reaches 9. How you specify the task is a bigger lever than the architecture.
Per-step success: demonstration prompt vs language prompt, percent
Both arms of the study used identical data, architecture and compute, so the gap comes from the prompt, not the model. Language wins small – 53 against 43 on seen tasks at a thousand hours – then loses decisively once pre-training grows, with post-training only overtaking a single in-context demo at two thousand episodes. The buyer’s version of this question is blunt: how long does it take your own people to point the machine at a new job. Three hours today. It needs to be five minutes.
Sources: Skild AI, “Introducing S1”, 18 August 2026; Patrick Schwarzkopf, VDMA, Stuttgart. Figures are company-reported.
S1 executing a task it has never seen, from one demonstration
What’s the one question your supplier would rather you didn’t ask?
Miles – or hours – between human interventions is the reliability measure frontier operators actually run their fleets on. And Europe is now the only region building an independent test regime for it.
Below ten miles between faults, someone walks alongside. Between ten and twenty-five, it needs remote supervision. Fraunhofer IPA tests collision force, cybersecurity, energy use and cleanroom particle class across 66 checks, and it’s taking the results into ISO Working Group 12. One more finding worth sitting with: Rhoda found two checkpoints from the same training run 28 points apart on the robot, and validation loss picked almost the worst of the two. Nothing short of a real trial tells you what you actually have. So ask it – what’s your mean time between human interventions, measured at a customer site, not in your lab.
Sources: Burro, Actuate 26; NIST humanoid benchmark proposal via The Robot Report, May 2026; Fraunhofer IPA, Stuttgart, September 2026; Rhoda AI, 10 September 2026.
Unitree closed 460 percent up, then fell 44. What was actually priced?
The first public price in this sector arrived and gave back 44 percent of its debut value – in the same quarter growth decelerated from 333 percent to about 40.
Adjusted profit fell 52.6 percent in the first quarter. Agility is going public by SPAC at a 2.5 billion-dollar pre-money valuation, with 65,000 operating hours and more than 300 million dollars of multi-year orders behind it – which is a different kind of number than a share price. The count of limited partners behind venture funds has halved since 2022, concentrating capital into fewer, larger cheques, while Cartwheel Robotics sits in involuntary Chapter 7. Don’t predict a date. Say instead what has to be true for the correction to stop.
Sources: Xueqiu.com, 14 September 2026; The Economist, 23 August 2026; Caixin; Bloomberg; Reuters, August 2026; Business Times, 24 August 2026; Agility Robotics and Churchill Capital Corp XI, June 2026; Robots & Startups, May 2026.
What does this industry look like in 2032, once the robots are ordinary?
Three of the five layers in this value chain don’t yet exist as businesses. The integrators who deploy the fleets and own the operating data will carry more weight than the makers.
Think of the bicycle. The Netherlands turned it into infrastructure, and the infrastructure was never the factory – it was the repair shop on the corner. So: how many bike shops does this country have, and how many robot workshops?
Five layers of the value chain, and how much of each exists as a business today
| Layer | What it is | Status |
|---|---|---|
| Components | One supplier covers 60–70% of all humanoid makers | Mature |
| OEMs | 150–200 brands on a much smaller number of real manufacturers | Consolidating |
| Integrators & fleet operators | Buy, deploy, lease, and own the operating data | Forming now |
| Second-hand & residual value | No resale market, no residual curve – which is why leasing beats buying on cost | Doesn’t exist |
| Service, parts & aftermarket | Workshops, limb swaps, batteries, wear parts | Doesn’t exist |
Agility, Apptronik and AGIBOT all sell robots as a service themselves today – which is what happens when the integrator layer hasn’t been built. It doesn’t survive the first fleet of a thousand machines. Any limb on the latest Atlas can be swapped in under five minutes; Renault’s motors overheat after two to three hours of heavy lifting. Both are workshop arguments.
Sources: Werner Kraus, Fraunhofer IPA, from fourteen company visits in China, July 2026; Aya Durbin, Boston Dynamics; Renault and Wandercraft, Stuttgart, September 2026.
What can Europe actually own in 2030?
China controls roughly 90 percent of magnet processing, and actuation is 40 to 60 percent of the bill of materials. That leaves Europe the integration, certification and operations layer.
Estimated bill of materials for Optimus Gen 2, USD thousand
Where the value sits along the chain
The hostile answer first: the component race is lost. Actuation alone is 40 to 60 percent of the total, and China holds about 90 percent of magnet processing. The constructive one: McKinsey expects the supply chain to split in two rather than one side winning, with Europe differentiating on safety-certified, high-assurance deployment. Export markets may accept Chinese hardware while restricting software and data flows – the mechanism behind both the FCC determination and the new Agibot plant in Serbia. Wandercraft is putting 350 units into Renault. Verity won the IERA award at ICRA in Vienna. The operations layer is open, and that’s where the recurring revenue sits.
Sources: McKinsey, April 2026; IDTechEx; South China Morning Post; FCC national security determination, August 2026; Anadolu Agency, 29 August 2026.
Ask for hours in paid operation before units shipped
Almost every published figure counts deliveries. The number that predicts value is time in production at a customer site.
Price the cost of teaching one task before you price the machine
Four hours, ten hours or six months is the difference between a pilot and a plant – and it never appears in a quotation.
Build the position in the three layers that don’t exist yet
Fleet operations, residual value and service. The component race is decided; the workshop on the corner is not.
