On May 28, the startup Shift announced that it would clean any New York City apartment for free. In a launch video, cheerful young men scrubbed toilets, vacuumed floors, and wiped down counters.
The catch? Cleaners wore baseball caps with cameras mounted under the brims. The company planned to record the workers’ actions and sell the data to robotics companies.
While the whole deal might have been a gimmick — the scheduling website notes in an FAQ that the offer is only available for a “limited time” — it’s still a perfect encapsulation of one of the most important trends in robotics today.
Early LLMs were famously trained to “predict the next word” across billions of tokens of text scraped from the Internet. Most roboticists expect we’ll need something similar to train general-purpose robots: an Internet-scale database of everyday tasks that robots can learn from.
But right now, humanity doesn’t have anything like that. The largest openly available dataset of robots performing tasks, ABC-130K, only has 3,500 hours of task demonstrations.









