Skip to content
AI Primer
release

Figure launches Index robot-training dataset with 16M video uploads

Figure says Index has collected 16 million robot-training video uploads from contributors in 108 countries. The company reports 264,000 app downloads and says contributors upload more than 30 minutes of video each second.

4 min read
Figure launches Index robot-training dataset with 16M video uploads
Figure launches Index robot-training dataset with 16M video uploads

TL;DR

The official Index post says the app is already uploading the equivalent of 4.9 years of human work each day. The same write-up describes a creator marketplace for home and workplace tasks, while an independent comparison from Encord lays out the gap between raw video and training-ready robot trajectories.

The creator app

Figure spent four months building a Figure-exclusive data pipeline before bringing Index out of stealth. The app launched on Google Play and the App Store, with people recording tasks in their homes and workplaces or booking creators to perform work on site.

Figure's immediate deployment context is physical work. One customer task required a robot to climb a ladder to a mezzanine and return at the end of a shift adcock_brett's mezzanine example, and Figure had described the coming update as critical to solving general robotics adcock_brett's teaser.

Diversity metrics

The country count is only the broadest measure of Index's coverage. Figure's public breakdown reports the following per 1,000 hours collected:

  • 373 unique tasks, spanning cooking, cleaning, laundry, logistics, restaurants, factories, and offices.
  • 1,146 unique manipulated objects, including the long tail of household and workplace items.
  • 116 unique environments, with variation contributed by each creator's home, workplace, and surroundings.

The examples include cleaning kitty litter, changing oil, bussing restaurant tables, making beds, folding laundry, and stocking retail shelves. Figure says the app accepts contributions from individuals and businesses, which gives the collection process access to settings a lab-built dataset would have to enumerate in advance, as described in adcock_brett's Index write-up.

The robotics data gap

Robotics has no equivalent of the text archive that helped accelerate language models. People rarely document ordinary physical interactions, and they do not routinely record how they do laundry or navigate an unfamiliar room, sarahookr's laundry example argued.

Code created a cleaner learning loop because models can compile programs, run tests, and verify outcomes sarahookr's code comparison. Physical tasks often have no clear intermediate reward, and reaching the same goal can involve many valid paths sarahookr's robotics caveat.

Index's approach addresses the missing raw experience with human-recorded video, but the format leaves an important systems question open. Figure's post describes video uploads and hierarchical text captions, while Encord's data-collection comparison describes robot teleoperation data as synchronized camera, joint, and force streams that still require structured annotation before training.

Synthetic data may reduce the gap, according to sarahookr's optimism, but sarahookr also described real-world collection as a capital-intensive effort required to cover enough distributions and the long tail sarahookr's capital warning.

The five-stage pipeline

Processing 30 minutes of uploads per second required Figure to rebuild its infrastructure around continuous availability, large-scale processing, and feedback to creators. The disclosed pipeline has five stages:

  1. Filtering: automated checks for technical, visual, and semantic quality.
  2. Fraud review: human analysts audit samples at the user level for attempts to evade automated checks.
  3. Deduplication: video segments are embedded and compared with accepted data; segments above a similarity threshold are discarded.
  4. Rebalancing: task quotas and embedding-based clusters preserve coverage beyond the task labels.
  5. Annotation: the system generates hierarchical text captions for each episode.

Figure says internal generalization results are already validating its data thesis, but the launch post gives no scores or task-level comparisons. It also does not announce a downloadable corpus, only the app and the collection system that produces the training data.

Robots as a service

Creators have earned $15 million to date. Figure's model lets people record their own daily work, while customers can book a creator through the app to handle chores at home or send one to a business, according to the Index launch post.

Figure describes Index as groundwork for ordering robots as a service, with human contributors supplying the demonstrations needed before that service can be automated adcock_brett's robot-service post.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 4 threads
TL;DR2 posts
The creator app2 posts
The robotics data gap5 posts
The five-stage pipeline1 post
Share on X