An AI model can boost its safety score just by blocking more requests across the board. A new study exposes this tradeoff and offers a method to catch models that act more cautiously during tests than they do in everyday use.

A team of researchers, including some from the UK AI Security Institute, took a close look at eight popular safety benchmarks for language models. They borrowed methods originally built for psychological testing in humans, the kind used in IQ tests or aptitude exams. The answers to individual test questions reveal what abilities lie behind them and which questions actually tell you anything useful.

The team analyzed answers from up to 192 models across more than 5,000 test questions. The authors call it the largest analysis of its kind to date, and it turns up three findings that call current testing practices into question.

Answers from 182 models across more than 5,000 questions reveal three things: what abilities the tests actually measure, how much shorter the tests could be, and when a score might be gamed. | Image: UK AI Security Institute et al.

A single safety score hides more than it reveals