Figure unveils Helix 2.5 after tests in 30 unseen homes

Figure unveils Helix 2.5 after tests in 30 unseen homes

Figure has introduced Helix 2.5, a humanoid control policy that it says completed bed making, towel folding and living room tidying across 30 previously unseen Bay Area homes. The company reported a 56% overall success rate, compared with 9% for an otherwise equivalent policy trained without its Index human behavior dataset.

The result is zero-shot with an important qualification. The homes, layouts and manipulated objects were unseen, but the three tasks were specified using fine tuning data collected elsewhere. Figure used a fixed checkpoint for each task across the evaluation homes, with no local data collection, weight updates or adaptation after deployment.

Key facts

  • Evaluation: 30 unseen homes in the Bay Area
  • Tasks: Bed making, towel folding and living room tidying
  • Helix 2.5 success: 56%
  • Policy without Index pretraining: 9%
  • Total attempts: 420, according to Startup Fortune

Strict completion criteria expose uneven performance

Figure awarded no partial credit. A successful tidying trial required the robot to place all 13 to 15 toys in a basket. Towel folding required every towel to be folded and deposited, while bed making required both pillows and the comforter corners to reach the top of the bed, with the comforter pulled smooth.

Startup Fortune reported that the robot completed 237 of 420 attempts. Bed making was the strongest task at 67%, or 94 of 140 attempts. Towel folding reached 62%, or 87 of 140, while toy tidying scored 40%. The spread suggests that the aggregate result should not be read as uniform household competence.

Evaluation objects were separated before testing, according to Figure, and checked against the task specification dataset using an AI model followed by human review. Initial object configurations varied between trials. Timeouts applied to individual objects or task elements, and any rollout requiring human intervention for safety was recorded as a failure.

The physical scope is broader than a tabletop manipulation test. The behaviors combine walking, stance adjustment, bimanual manipulation and perception in furnished spaces. Figure also showed qualitative examples of the robot stepping back, repositioning and walking around a bed to recover from errors, although it did not provide a separate recovery rate.

Index pretraining accounts for most of the measured gain

The clearest experiment compares two policies trained on identical task data. Architecture, optimization, hyperparameters and evaluation were held fixed. One policy began with random weights, while the other used Helix 2.5 weights pretrained on Index, Figure’s dataset of human behavior.

The randomly initialized version succeeded in 9% of zero-shot trials, against 56% for the Index pretrained policy. Figure says no individual evaluation task represents more than 1.90% of the Index dataset, an effort to show that the pretraining material was broad rather than dominated by the three tested chores.

Figure also compared Helix 2.5 with a previous Helix 02 policy for the same task. Helix 2.5 reportedly matched its success rate with half the task specific adaptation data, while operating across the unseen homes. The company did not disclose an absolute amount of adaptation data in the article, so the claim remains a relative comparison with one representative Helix 02 behavior.

Scaling result measures action prediction, not task completion

Figure trained four models on nested Index subsets spanning an eightfold increase in pretraining data, while holding model size and downstream training constant. Held out robot action prediction loss declined with each doubling. Using the smaller runs, the company says it forecast the largest model’s loss to four decimal places, with an error equal to 0.54% of the variation across the full data range.

This is evidence for predictable human to humanoid transfer under the tested setup, but the metric is action prediction loss rather than autonomous household task success. Establishing the same scaling relationship on completion rates across more behaviors and environments would be a harder test.

Figure says Index is adding roughly 35 minutes of human experience each second, and the company has committed $3.5 billion of compute to Helix training. Helix 2.5 gives that scaling strategy a substantial company run benchmark, while its 56% completion rate leaves considerable room between generalization and dependable household operation.

Sources: FigureAI, Startup Fortune

Similar Posts

New Report

The Humanoid Robot Supply Chain

Supplier Strategy and Market Positioning 2026–2027

Get the Report

New! 2026 Humanoid
Robot Market Report

198 pages of exclusive insight from global robotics experts — uncover funding trends, technology challenges, leading manufacturers, supply chain shifts, and surveys and forecasts on future humanoid applications.

Aaron Saunders
Featuring insights from Aaron Saunders, Former CTO of Boston Dynamics,
now Google DeepMind
Get the Report
New Report

The Humanoid Actuation Report

Fundamentals, Architectures, Vendors and the Ten-Year Outlook

Get the Report
New Report

Humanoid Foundation Models

The brains are being rebuilt

Get the Report