Humanoid robot company Figure is launching a gig economy platform called Index to gather robot training data from human workers wearing cameras and sensors.John KoetsierHumanoid robot company Figure has launched Index, a gig economy marketplace where you can hire human workers to do jobs while wearing cameras. The goal: massively increasing data capture to train robots who will eventually do those jobs.CEO Brett Adcock says the company has already paid $15 million to people filming themselves do housework and gig work for companies. Index is currently processing 30 minutes of video uploads every single second, the company says, and per thousand hours, Index captures data on 373 different tasks, 1,146 unique objects, and 116 environments. The Index app already has 264,000 downloads with 44,000 weekly active users, resulting in 16 million videos uploaded from 108 countries to Figure’s data centers.The company calls its gig workers “creators,” and says each one of them “brings an unseen environment, unfamiliar objects, and their own idiosyncratic way of completing a task, the kind of long-tail variation that's nearly impossible to define upfront.”All of the video processing is expensive: Figure expects to spend a billion dollars over the next 12 months on data and compute, Adcock added on X. Figure is making three big bets hereFigure’s billion dollars is making three big bets.MORE FOR YOUThe first bet is that volume and diversity are the whole ballgame. Get enough footage of enough different people doing enough different things in enough different rooms, and you get generalizable training advantage. That’s plausible, but nobody has firmly established robotics scaling laws the way they exist for language. In text, more tokens fairly reliably buys more capability. Will the same happen in manual manipulation?We’ll see. The second bet is that passive video is enough. A phone recording of you folding a shirt explains nothing about joint angles, gripper state, force or torque, how hard you pinched the fabric … everything the robot actually needs to command its own actuators has to be inferred. Figure’s Project Go-Big shows human-to-robot transfer from egocentric video shot in homes, so this is plausible to an extent. But we’ll have to see how it plays out in scale.The third bet is that you cannot generate your way out of this. That synthetic data is not the answer.The data "has to come from the real world: a global sampling of physics captured across every environment on earth," Figure says. That is a sampling argument. Figure is essentially saying that the world contains more physical weirdness than any simulator will think to produce, and the only way to capture it is to do it IRL. There are opposing views, of course. Nvidia’s physical-AI stack is a bet in the other direction — that you can take a small real dataset and multiply it to sort of digitally feed the 5,000. For GR00T N1, its open humanoid foundation model, Nvidia’s researchers collected 88 hours of in-house teleoperation on a Fourier GR-1, then fine-tuned video generation models on that footage to produce "827 hours of video data," thereby augmenting the original "by around 10x." Then they separately generated 780,000 simulation trajectories, which the paper pegs as "equivalent to 6,500 hours."In other words, roughly 1% of GR00T N1’s corpus was captured from reality. The other 99% came out of a GPU.This is a cheaper approach. Quicker too.But Figure is spending a billion dollars to buy the real hours of real work by real humans.Adcock hasn't always drawn the line this hard. Back in 2023 he listed the options evenhandedly: "We can do it synthetically through simulation, we can do it through human demonstrations. We can do it through the robot itself doing actions and learning from those actions if things went well or not." Three years and a $39 billion valuation later, the money says which one he picked.World’s biggest dataset?The numbers here are impressive.Sixteen million videos, 30 minutes of upload per second, 4.9 years of human work every day … if this is sustainable, Index passes every public egocentric dataset in a week.For scale, Meta’s Ego4D, the reference egocentric video corpus, is 3,670 hours. Ego-Exo4D is roughly 1,286. Apple’s EgoDex, with paired 3D hand pose, is 829. On the robot side, Open X-Embodiment pooled more than a million trajectories from 60 datasets across 34 labs, and AgiBot World Colosseo logged about 2,976 hours from 100 dual-arm robots. Generalist AI says its GEN-0 corpus holds 270,000 hours of real dexterous manipulation and grows by 10,000 hours a week.There are some caveats. Figure has released an ingest rate, not a corpus size, and Figure’s own five-stage pipeline will throw a lot of it away. An hour of phone video is not an hour of robot data: it’s just simply not always usable. In fact, 10,000 hours of teleoperated robot data with action labels might be worth more than a million hours of somewhat-more-random video. Teleoperated data costs $50 to $200 an hour, Flikforge CEO Jeff Allen estimates, while egocentric human capture trades at $2 to $5. That price gap is an estimate of the value gap.If Figure’s bet pays off, however, it gets potentially the best-trained robots in the world.Figure and the home market: humanoids in the homeThis announcement is a pretty big tell that Figure is explicitly aiming at the consumer market: home robots.Look at what Figure is actually paying to have filmed: making beds, folding laundry, cooking, cleaning, watering plants, loading groceries, taking out the trash. There’s warehouse and retail and restaurant work in the mix too, but the center of gravity is pretty domestic.The diversity metric Figure published is 116 unique environments per 1,000 hours collected. That’s not a factory or a warehouse. It might be restaurants and retail, but it’s a lot of very individual homes. A house is a different environment pretty much every single time, which is exactly why homes are the hardest market in robotics.A training set is a roadmap: you collect data for the product you intend to ship. And Figure is collecting millions of videos of people’s kitchens and bedrooms.Being a creator must be surrealIt must be surreal to be a “creator” in the Index gig economy: the work you are doing is being done precisely so that humans will not be doing it in the future. That has to come with a bit of a wince.On the other hand, as a professional word-slinger, I guess I know exactly how that feels.Ultimately, the hope is, we create a world in which truly human work remains, and people are still able to earn a living while the robots and the software do much of the other work.We will, as they say, see.
Figure’s Billion-Dollar Bet: Gig Platform For Humans To Generate Robot Training Data
Humanoid robot company Figure will spend a billion dollars to get robot training data from humans doing gig work in homes and businesses.







