For the last few years, AI engineering has operated under a surprisingly simple assumption:

Bigger models are better models.

More parameters. More training data. More GPUs. More compute.

And to be fair, that strategy has worked remarkably well. Scaling has produced huge improvements in language understanding, coding, reasoning, multimodal capabilities, and general-purpose AI.

But engineering is rarely about maximizing one metric.