OpenAI’s Astra just posted a 169 on the Epoch Capabilities Index, the composite benchmark maintained by independent research organization Epoch AI. That score aggregates performance across more than 37 individual evaluations spanning math, coding, science, and agentic tasks, and it represents the highest number any AI model has achieved on the index.
To put that in perspective: Claude 3.5 Sonnet sits at 130 on the same scale, and GPT-5 scored 150.
What the numbers actually look like
The raw benchmark results behind that composite score are striking on their own. Astra recorded a perfect 100% on ExploitBench, a cybersecurity evaluation that tests a model’s ability to identify and exploit software vulnerabilities. It hit 99.9% on ARC-AGI-3 when run through a Provider Adapter harness, a test designed to measure abstract reasoning and pattern recognition. And it scored 98% on FrontierMath Tier 4, the most demanding tier of a mathematics benchmark built to challenge frontier-level models.
Astra also posted strong results in math, puzzle-solving, and coding suites, though the specific scores on those individual evaluations weren’t broken out in the same detail.












