OpenAI is urging a broader view of frontier-model evaluation: benchmark results reflect not only the...

OpenAI is urging researchers, evaluators, and AI buyers to treat benchmark results as measurements of...

OpenAI is urging a broader view of frontier-model evaluation: benchmark results reflect not only the...