The old ways of testing and evaluating new frontier AI models need a rewrite.

The old ways of testing and evaluating new frontier AI models need a rewrite.

Frontier AI models are outrunning the benchmarks built to measure their hacking skills, just as a 1 August US deadline for new testing standards looms.