Robot learning is getting better very quickly.

We now have better policy architectures, more capable simulation environments, standardized dataset formats such as LeRobot, and an increasing number of public robot manipulation datasets.

But there is a basic question that is surprisingly difficult to answer:

How do we know whether a robot training dataset is actually usable?

A dataset can be perfectly readable and still be a poor training dataset.