Chinchilla is quoted as a rule about how many tokens to train on. That is the part that stopped being followed. The reason it stopped is a different piece of arithmetic that nobody published as a headline, and it is short enough to do here.
What a scaling law is
An empirical relationship between resources and pretraining loss. Across many training runs at different sizes, loss falls as a power law in parameters, in data and in compute — meaning that on log-log axes the points fall on a line, and that each constant multiple of a resource buys a constant subtraction from the loss.
Two consequences are worth internalising before any specific result. Diminishing returns are built in: the step from 1B to 10B parameters and the step from 10B to 100B buy the same amount of loss. And the quantity being predicted is loss on held-out text, which is not the same as anything you care about — a point the emergence debate is largely about.
Kaplan, then Chinchilla






