Figure's Helix 2.5 Cleans 30 Strangers' Homes It Has Never Seen

Figure's new humanoid policy walked into 30 unseen Bay Area homes and tidied, folded, and made beds with zero prior data from those spaces.

·
·
Read5 min
TypeNews
  • Figure released Helix 2.5, a humanoid policy tested zero-shot in 30 unseen Bay Area homes.
  • One foundation model handles tidying, towel folding, and bed making with whole-body control.
  • Index pretraining lifted zero-shot success from 9% to 56%, a 6x gain over from-scratch training.
  • Helix 2.5 matched Helix 02's success rate with half the task-specific data, across 30x more environments.
  • First human-to-humanoid scaling law: largest run's loss forecast to four decimal places before training.
  • Figure has committed $3.5B of compute to Helix and is scaling Index rapidly.

Helix 2.5 enters 30 unfamiliar homes

Figure has announced Helix 2.5, reporting that its humanoid robot performed household tasks in 30 Bay Area homes excluded from its training and adaptation data. The company rented the homes and tested the robot on tidying living rooms, folding towels, and making beds using each property’s furniture and work surfaces.

Figure uses “zero-shot” to mean that the robot received no data collection, fine-tuning, or adaptation involving the test homes or their objects. The three behaviors were still specified and trained with demonstrations collected elsewhere. Figure says this is the first demonstration of zero-shot, whole-body generalization across this many homes on a humanoid.

Each behavior is long-horizon, meaning success depends on completing a sequence of actions without allowing an early error to derail the task. Figure awarded no partial credit.

Behavior End-to-end completion criterion
Tidy a living room Place every target toy in a basket
Fold towels Fold and place every towel
Make a bed Position the comforter with its corners reaching the top third of the bed

Thirty homes test the deployment gap

Robot-learning systems often lose reliability when moved beyond the spaces used to collect their training data. Tabletop arms operate within fixed workspaces, while wheeled robots need enough clear floor to turn and approach objects. Homes vary in room dimensions, furniture placement, lighting, clutter, fabrics, and available walking space.

Helix 2.5 treats those variations as a whole-body control problem. The robot may need to walk until an object becomes visible, adjust its stance before reaching, coordinate both hands, and move its head or torso to improve its view. Figure adapted one broadly pretrained foundation model into three behaviors spanning locomotion, rigid-object handling, fabric manipulation, bimanual coordination, and active perception.

Index lifts success by 47 points

Figure isolated the effect of pretraining by adapting two policies with identical task data. One began with random weights. The other began with Index, Figure’s large-scale collection of human-behavior video.

Zero-shot success rates for policies trained with and without Index pretraining
Figure’s reported zero-shot success rates with and without Index pretraining.

In evaluations that Figure describes as blind, the policy trained from scratch completed 9% of trials. The Index-pretrained policy completed 56%, an increase of 47 percentage points and roughly 6.2 times the baseline rate. A successful trial required finishing the entire assigned task.

Helix 2.5’s base model began from random initialization and was pretrained entirely on Index. Helix 02 used a pretrained vision-language model as its starting point. The newer training setup ties the measured improvement more directly to human-video pretraining.

Less adaptation data, wider deployment

Figure also reports that Helix 2.5 matched the success rate of an earlier Helix 02 policy while using half as much task-specific adaptation data. The resulting behaviors were then evaluated across 30 unseen homes without further adjustment.

The two figures describe separate dimensions of the result: a twofold reduction in adaptation data and testing across 30 deployment environments. They do not combine into a single multiplier for overall model quality.

Figure’s demonstrations show the robot stepping backward to improve its reach, changing stance after a poor approach, and walking around a bed to repair a fold. These recovery actions address a common failure mode in long tasks, where one missed grasp or awkward position can corrupt every subsequent step.

Figure trained four models on nested subsets of Index spanning an eightfold range of pretraining data. Model size and downstream task training remained fixed, allowing the company to measure how additional human video affected held-out action-prediction loss.

Scaling curve for transfer from human video to humanoid robot action prediction
Held-out action-prediction loss declined as Index pretraining data increased.

Action-prediction loss measures the gap between the actions a model predicts and the recorded actions in held-out examples. Lower loss indicates better prediction, though the metric serves as a proxy for physical task performance.

Loss declined predictably with each doubling of Index data. Using the smaller runs, Figure says it forecast the largest run’s test loss to four decimal places before training began. The forecast error equaled 0.54% of the loss variation measured across the full data range. Figure describes the result as the first human-to-robot transfer scaling law measured on a humanoid.

A stable scaling relationship would help robotics teams estimate the likely return from additional data and compute before committing to expensive training runs. Figure says it has allocated $3.5 billion of compute to Helix and that Index now collects about 35 minutes of human experience video every second.

The evidence has firm boundaries

  • Task scope: The evaluation covers three predefined behaviors with explicit completion rules. It does not demonstrate open-ended household assistance.
  • Reliability: The pretrained policy failed 44% of pooled end-to-end trials.
  • Reporting: A pooled success rate can conceal variation among tasks and homes. Per-task results, trial counts, and uncertainty estimates would clarify statistical strength.
  • Scaling evidence: The curve comes from four models across an eightfold data range and measures prediction loss rather than physical completion rates.
  • Validation: Figure reports its own evaluation, and no independent replication is cited.
  • Availability: Figure has announced no public model weights, API, price, or release date.

A concrete recipe for the next experiments

For robot-learning teams, Figure’s reported ablation supports a specific development strategy: pretrain on broad human-behavior video, adapt with smaller robot datasets, and evaluate complete tasks across many physical sites. The 47-point gap suggests that pretraining can improve recovery and generalization when deployment environments vary.

Broader task coverage, trial-level reporting, independent replication, and a demonstrated link between action-prediction loss and real-world completion rates will determine how far the approach transfers. Helix 2.5 provides a measurable starting point for testing those questions.

Trending
  • No trending articles

Comments

avatar

Next Reads