Figure says Helix 2.5 raises zero-shot household-task success from 9% to 56% across 30 unseen homes
Figure reports that Helix 2.5 raised zero-shot household-task success from 9% to 56% across 30 unseen rental homes without retraining. The company released four hours of video showing the humanoid performing the tasks.

TL;DR
- Figure says Index pretraining lifted Helix 2.5's full-task success from 9% to 56% in 30 unseen homes, according to rohanpaul_ai’s result post.
- The evaluation used 30 Bay Area rentals with no local data collection or retraining, which adcock_brett’s 30-home post describes as the test for real generalization.
- Index data followed a smooth 1x-to-8x action-prediction-loss curve, with rohanpaul_ai’s analysis noting that this forecasts model loss rather than household-task reliability.
- Figure released nearly four hours of robot footage via adcock_brett’s four-hour reel, rather than only a short highlight video.
Figure’s evaluation rubric gives each toy a one-minute timeout and each towel three minutes, while any safety intervention aborts a run as a failure. Its earlier Index announcement said the human-behavior collection app had passed 16 million uploaded videos from more than 44,000 weekly active creators.
30 homes, three whole-body tasks
Figure defines zero-shot narrowly: it collected no data, performed no fine-tuning, and made no adaptation in the evaluation homes, while evaluation toys, towels, and bedding were excluded from task-specification data. The technical report says an AI model followed by human review checked that held-out-object condition.
The fixed behaviors were:
- Living-room tidying: gather 13 to 15 scattered toys into a basket.
- Towel folding: fold every towel and place it in a basket.
- Bed making: move both pillows and the comforter corners to the top of the bed, then smooth the comforter.
Each task used one fixed checkpoint in all 30 homes, as Figure’s evaluation setup specifies. The jobs combine navigation, active perception, bimanual manipulation, and posture control around clutter and unfamiliar furniture.
The pretraining control
Figure trained two policies on the same task-specification data: one from random weights, and one initialized from Helix 2.5’s Index-pretrained weights. Its controlled comparison held architecture, optimizer, hyperparameters, downstream data, and evaluation fixed, leaving pretraining as the stated experimental variable.
Success meant completing every required subtask, with no partial credit. Figure’s rubric required every toy to reach the basket, every towel to be folded and placed, and the full bed configuration to meet its criteria.
The 56% figure represents 237 successes in 420 trials, according to Humanoids Daily’s count. The same report also separates environmental zero-shot transfer from task learning: the three household behaviors were still specified with training data gathered elsewhere.
Figure separately says Helix 2.5 matched a comparable Helix 02 behavior using half the adaptation data, while generalizing that behavior across the 30 evaluation homes in its Helix comparison.
Index scaling curve
Figure trained four models on nested Index subsets spanning an 8x increase, while keeping model size and downstream training fixed. Its scaling-law report says smaller runs forecast the largest run’s held-out action-prediction loss to four decimal places, with error equal to 0.54% of the variation across the full range.
The curve measures next robot-action prediction, not end-to-end household completion. As rohanpaul_ai noted, predictable loss reduction does not yet establish an equally predictable scaling law for the 56% task-success result.
Four hours of robot runs
Figure also published its long-form evaluation footage alongside the technical post.
The “breakthrough” teaser
Before the technical disclosure, rohanpaul_ai flagged an incoming Figure reveal and koltregaskes called the scheduled demo a breakthrough.
The label also drew skepticism. koltregaskes’ follow-up said Figure had used “breakthrough” three or four times before, including at least one reveal they considered minor.