On ulam.ai's ErdosBench, Astra took first place. The benchmark covers 226 open math problems inspired by the famous Erdős problems. Astra scored 3.23 and solved 106 problems, 43 of them completely. It disproved 27 more. It also disproved 27 others.
Compared to Sol, which solved 78 problems at maximum reasoning, Astra showed stronger scientific writing and was less prone to overblown claims. In some cases, it actually understated its own results. Overall, benchmark developer Przemek Chojecki called it "a solid 5%-10% gain on various math-research skills tested," but the benchmark is far from saturated.
GPT-6 Astra leads the ErdosBench with a score of 3.23 and 106 solved problems, followed by GPT-5.6 Sol at 3.12 and 78 solved problems. | Image: ulam.ai
OpenAI chose not to optimize for math
OpenAI could have made Astra much stronger in math research but decided against it, even though the company had put math wins front and center in its first announcement. In his essay "An Alien Mind," OpenAI chief scientist Jakub Pachocki writes, "[…] we believe we could make the models better at specifically mathematics research with additional focus, but we do not prioritize this direction because of the urgency we feel about RSI and automated alignment research, as I will discuss later."













