A model scores 90% on MMLU. It scores 95% on GSM8K. It scores 98% on HumanEval. The press release says: "State-of-the-art. Approaching AGI." You try the model yourself. It's good. It's not that good. It makes mistakes on simple tasks. It struggles with reasoning. It fails on tasks that are not in the training data. The benchmarks are lying. They are broken.

This is the problem with standard benchmarks. They are saturated. They are leaky. They reward memorization. They do not measure true capability.

The Problem with MMLU

MMLU (Massive Multitask Language Understanding) is a multiple-choice benchmark.

The Concept: