OpenAI is urging a broader view of frontier-model evaluation: benchmark results reflect not only the model being tested, but also the surrounding system used to test it. In its official playbook for trustworthy third-party evaluations, the company says API settings, prompting, tool access, state management, compute budgets, scoring, and harness design can materially affect conclusions about model capability and safety.
That framing matters as evaluations increasingly assess agentic, tool-using systems rather than isolated text generation. A model may perform differently when it can use tools, retain or compact state, receive a different prompt, or operate under another budget. OpenAI’s central argument is that evaluators should make those choices visible, test their validity, and calibrate claims to what an evaluation actually measures.
Why the harness is part of the result
A harness is the evaluation environment around a model. It can include the prompts and instructions supplied to the model, the tools it can call, the way its state is managed, limits on computation or attempts, and the mechanism used to score its output. These choices are not simply implementation details when they influence the behavior an evaluator observes.









