Figure 1. Art courtesy of Dr Singularity. An explosion of progress in mathematics may be upon us, thanks to next-generation AI.One clear sign of AI progress is the shift in what constitutes news. A year ago, a Gold medal at the Math Olympiad was news. Now, math benchmarks are saturated, and this weekend, OpenAI announced that their next-generation frontier model Astra solved ten major open math problems.OpenAI reported that an internal version of its forthcoming Astra model generated solutions or substantial progress on ten long-standing problems in mathematics. They showed their work; each proof solution was subsequently prepared by humans and formally verified with Lean certificates:We’re releasing 10 such Astra proofs, complete with lean certificates and CoT walkthroughs for each of them.The 10 results are published in a lengthy 249-page paper, and cover a range of topics in mathematics, across sphere packing, coding theory, group theory, circuit complexity, quantum games, lattice cryptography and extremal combinatorics:Improving a 1978-era lower bound on high-dimensional sphere packing.An explicit construction of a non-sofic group, open since Gromov’s 1999 introduction of soficity.A disproof of Connes’s rigidity conjecture that certain groups are uniquely determined by their von Neumann algebras.Resolutions of several Erdős problems: #146, #180 in extremal graph theory; and #183, proving a super-exponential lower bound for multicolor triangle Ramsey numbers.Resolving Ehrhart’s volume conjecture: Determining, in every dimension, a bound on the maximum possible volume of a convex body whose centroid is its only interior lattice point.A result on quantum parallel repetition, relevant to post-quantum cryptography.A new lower bound on algorithmic circuit complexity, and closest-vector-problem hardness.Don’t worry if you don’t understand the problems or the jargon; most of us non-mathematicians don’t either. Nor can we attempt to understand or verify the paper or proofs; the professional mathematicians are checking this work using LEAN. An example is the proof that The Erdős–Simonovits degeneracy conjecture fails at every level.OpenAI got these results with Astra, their next-generation AI model that may become GPT-6 and/or a Mythos / Fable 5 equivalent AI model. They claimed the cost was only $2000 in Sol equivalent tokens, a low compute cost for 10 meaningful novel math results. However, this doesn’t include the human-level preparation plus independent verifiability via Lean, and OpenAI’s report left open a few questions: How much human guidance was there? Who decided which problems to pursue? Did OpenAI cut the AI loose on many more problems, but only reported successful proofs? Was it 1,000 problems and got 1% success rate with 99% failures, or was it more targeted?Unanswered questions and hype aside, it’s still a huge accomplishment for OpenAI and AI generally. It’s been noted that several proofs are counterexamples to long-held beliefs, which increases their value as it changes understanding in these domains.Astra’s ten math proofs will still require detailed scrutiny and validation by independent mathematicians, but it shows next-generation frontier AI models are moving beyond solving benchmark-style mathematics toward producing original, formally checkable research contributions.What was achieved here is what we can call Generative Math. In the same way AI image generation models create novel images, AI coding models craft new software, and AI music generation generates new songs, the Astra AI model generated new mathematical proofs, all of them complex, difficult, novel, and useful.Generative Math is not just ‘doing harder math’ as a brilliant PhD-level math student at a higher-level IMO, it’s AI engaged in mathematical invention. This marks a new milestone for AI mathematic reasoning.Invention is implicit in generative AI models. The very randomness that plagued early LLMs and led to problems such as hallucinations makes them just random enough to be inventive as they generate.However, randomness alone does nothing but stumble in the dark. It takes a keen amount of reasoning and domain-specific intuition to avoid wasting time and to productively explore a huge problem space. Math prodigies and chess grandmasters possess that skill of distilling an enormous search space into the essential mental model and reason with high focus on what matters.OpenAI put this new level of capability on their 5 stages of AI roadmap some time ago, putting “Innovators, AI that can aid in invention” on the stage beyond Agentic AI. We observed that the latest level of AI capability is long-horizon multi-agent AI, which is consistent with the long-horizon reasoning needed for AI invention. The two capabilities coincide.Even if OpenAI’s report doesn’t report on the human effort and scaffolding to make this happen, the plain fact is that these problems have remained open for decades because they are hard. As Pavel on X put it:I’ve spent well over 10,000 hours studying math in my life, yet I can’t understand these proofs, at least not without weeks of digging deep into each topic. What’s more, none of my math PhD friends know much about these problems either, and they can’t verify most of them without working directly in the field (yes, math is VERY diverse). LLMs are getting smarter than the experts themselves, and I’m not sure we have enough bright human minds to verify everything that will come out of them in the coming years. Remember when we compared AI intelligence to PhD students? I think we’re past that.If we are past asking if AI is as good as PhD student, how good IS this work? Thought experiment: Imagine if these results were the work record of a young math researcher. Ten profound results that solve long-standing problems in mathematics in a matter of months. Such a prodigy would be called an Einstein-level genius and be on track to win a Fields medal. That sounds like math superintelligence.Some mathematicians have stated that the results are genuinely significant and more advanced than prior AI math results, but they are not revolutionary in kind. AI can excel at induction/deduction within existing conceptual frameworks, but cannot yet perform the “abductive jump” of inventing new conceptual worlds:Current LLMs can fill the gaps of knowledge left by humans. But they can’t jump beyond the external boundary of existing knowledge.This is likely true. The AI has been trained on existing patterns of reasoning. It doesn’t invent new forms of reasoning; it’s real superpower is hammering with patterns of deep reasoning on a problem at a level that goes long past human endurance and patience. There’s a lot of space in the gaps of human knowledge where such dogged effort can yield results.We might back off and call this gap-filling Generative Math to be closing in on Math AGI, accomplishing useful human-level math research with AI. We can distinguish between types of human-level AGI this way:Math AGI: Capable of the generalized work of a professional mathematician.Software AGI: Capable of the generalized work of a professional software engineer.Virtual knowledge-work AGI: Capable of the work of a knowledge worker working in an online environment.Robotic AGI: Capable of the embodied activities of a human being or worker such as a nurse, bricklayer, chef, factory worker, or manual laborer.We are near Software AGI and Math AGI because the reward signal for RL training in the math and coding domains is clearest, and coding has been a lucrative area for AI labs to focus on.An AI skeptic might admit AI shines in math because it has a measurable and verifiable reward attached to each problem, while claiming that such verifiability defines AI’s limits. Gary Marcus is correct that “Expertise in one domain does not at all guarantee expertise in all or even most domains.” Progress in math does not automatically translate to progress in other areas of science that are less amenable to clear verifiable rewards. However, progress in generative math does translate into progress on rigorous deep reasoning, a useful skill of pruning extremely complex solution spaces to productively reason on it. If AI models crack that “rigorous deep reasoning” skill, it could be a key to superintelligence and will broaden its applicability across domains. Figure 2. Tracking the AI performance on AI benchmarks versus human performance show that we are asymptotically approaching a limit where AI performance matches human performance levels across many benchmarks.AI intelligence unlocks AI automation when it has a supporting AI ecosystem of memory, skills, agentic harnesses, tool use, and more. Generative Math with AI unlocks a new kind of mathematics development when it has the supporting infrastructure as well.The leading mathematician Terence Tao gave a recent talk on what this might look like. His views, prior to OpenAI’s announcement, was that frontier AI had gotten to a “junior human co-author” level of capability in math. As stated in this summary of Tao’s views on AI and mathematics:Tao doubts that genuine “artificial general intelligence” is within reach of current tools but argues a weaker and still very valuable “artificial general cleverness” is becoming real: the ability to solve broad classes of problems by ad hoc, often brute-force or stochastic means that are fallible and uninterpretable, yet succeed at non-trivial rates when coupled with strong verification. … today’s tools are best viewed as stochastic generators of sometimes-clever, often-useful outputs — which yields the characteristic “useful yet unsatisfying” feeling of a magic trick once explained.Tao paints a scenario of AI as an efficient complement to mathematicians in a ‘factory-style’ big mathematics. Just as Supercolliders use power and brute force to understand particle physics, mathematicians can harness AI to brute-force solve many aspects of mathematical understanding. The AI magnifies the mathematicians’ productivity but changes their efforts to manage and curate how the AI’s focus, akin to deciding what Supercollider experiments to run.OpenAI has shown that Astra has a deep domain understanding of mathematics that translated into significant novel findings. DeepMind has produced similar profound results in other areas of science. There will be more to come. AI will disrupt and accelerate many processes of research, discovery and invention in mathematics and science.However, it will be neither instant nor complete. This is not a ‘singularity’. Instead, mathematics and science research is becoming industrialized; AI is making science research cheaper and more scalable at a level akin to what early 1800s industrialization did to textiles. The acceleration this implies is enormous.AI will play ever larger role in researching, proving and verifying mathematical hypotheses, but like with industrial factories, man and machine will both play a role, leveraging capabilities with AI. Output could accelerate by an order-of-magnitude. In coming years, hand-crafted AI-free science and math may become an artisan activity.It is not the most significant advance in mathematical history, but this milestone is both important and shows a clear direction. AI for mathematics has gone from solving 8th grade level math in 2023, acing high school level math benchmarks in 2024, achieving Gold-level in IMO math in 2025, and now solving major novel mathematical open questions in 2026. We’ve come very far, very fast.OpenAI has shown with their 10 AI-generated proofs that AI for Generative Math is here to stay. Astra’s concrete, checkable advance in AI reasoning indicates that AGI-level AI for Math is on the horizon. This advance in AI will automate significant parts of mathematical research and accelerate mathematical progress tremendously.OpenAI has shown that their new Astra AI model is capable of levels of deep reasoning that could unlock not just mathematics, but other areas of science and engineering invention via deep reasoning. Whether and how this will spill over into other domains is unclear, but major AI labs have other areas of science on their radar. AI will eventually take them on, and the question won’t be if AI accelerates scientific progress, but how much and how soon will it do so.