OpenAI is urging researchers, evaluators, and AI buyers to treat benchmark results as measurements of a model in a particular setup, not as permanent measures of a model's standalone capability. Its new guidance on third-party evaluations says the harness, resource budget, available tools, and approach to memory and context management can materially change observed performance, especially on long-running, multi-step tasks.
In OpenAI's guidance on trustworthy third-party evaluations, the company calls for clearer reporting about the conditions behind a score. The central message is straightforward: when a model's results improve with more compute, better tool access, or a more capable agent harness, the finding should be described as performance under that harness and budget. It should not be presented as a fixed capability ceiling.
That distinction matters as frontier AI systems are increasingly evaluated as agents rather than as single-turn chat systems. An agent may need to retain useful information, choose and use tools, recover from failed steps, and manage its context across an extended task. In those cases, evaluation design is not a minor implementation detail. It is part of what determines the result.








