The scoring is genuinely new. The matrix underneath it is a solved problem from 2015.

Three questions teams actually ask about their LLM systems:

Is my classifier right?

Did my prompt change help?

Is the cheap model good enough?