OpenAI's Astra Solves 10 Long-Standing Math Problems
OpenAI says an internal, unreleased version of its next major model, Astra, has produced new results on 10 mathematics and computer science problems that had gone unsolved for at least a decade — and this time, unlike a similar claim last year, outside experts are backing it up.
What Astra actually did
On August 1, OpenAI published a 249-page manuscript collection along with Lean 4 certificates — machine-checkable proof files — for all 10 results, posted publicly on GitHub under an open license. The headline result is an explicit construction of a non-sofic group, a question left open since 1999. Astra also disproved a 1980 conjecture by mathematician Alain Connes, proved a separate result known as Ehrhart's volume conjecture, and resolved three problems from Paul Erdős's long-running catalog of open problems, including one on Ramsey numbers. The rest of the list spans high-dimensional sphere packing, coding theory, arithmetic circuit complexity, quantum computing, and lattice cryptography.
The Lean certificates matter because they remove guesswork: Lean's verification system either compiles a proof or it doesn't, so there's no room for a model to bluff its way through a step. OpenAI reported the certificates carry a "sorry" count of zero, meaning no step in any of the 10 proofs was left unproven. Human researchers turned Astra's reasoning into published manuscripts, though OpenAI says the underlying mathematical arguments came from the model itself. The company estimated the total compute cost for all 10 solutions at roughly $2,000 in API tokens.
Why this claim is landing differently
OpenAI has stumbled here before. In October 2025, then-VP of science Kevin Weil claimed GPT-5 had solved 10 previously unsolved Erdős problems — a claim Thomas Bloom, who maintains the erdosproblems.com database, called "a dramatic misrepresentation," since the model had mostly just found existing papers Bloom wasn't personally aware of. Weil deleted the post.
This time, Bloom reviewed the Astra results himself and called them "big news," rating them above an Erdős counterexample an internal OpenAI model produced in May that he also helped verify. OpenAI research scientist Noam Brown, associated with the company's test-time reasoning work, called the results "a major step for scientific reasoning." Still, none of the 10 results has been through formal peer review yet, and a Lean certificate confirms a proof is internally valid, not that it correctly captures what the original open problem was asking, or that it matters. That judgment still requires a mathematician.
The release also lands amid friction between AI labs and the math community. In June, the International Mathematical Union endorsed the Leiden Declaration, warning that AI companies are using published research without consent and bypassing peer review. Astra itself remains unreleased; OpenAI has not said whether it will ship as GPT-5.7, GPT-6, or under another name, and any launch would go through the federal AI safety review already applied to GPT-5.6.
Why this matters to small and medium businesses
This isn't really a math story — it's an early look at reasoning power headed for consumer and business AI tools. Astra is built to run long, multi-step tasks by coordinating multiple reasoning passes over extended periods. That same underlying capability, when it eventually reaches ChatGPT or the API, is what will make AI more useful for genuinely complex business work: multi-step research, long documents, tangled scheduling or logistics problems, not just quick one-off answers.
Don't take capability claims from any AI company at face value, including this one. OpenAI's own history here, an overstated claim in October followed by a more credible one in August, is a useful reminder to wait for independent verification before changing how you rely on a tool, especially for anything consequential to your business.
Cost is trending toward "cheap enough to not think about." OpenAI's estimate of roughly $2,000 in compute to crack 10 problems that stumped expert mathematicians for decades is a data point on how fast the price of serious AI reasoning is falling — worth keeping in mind if you've written off AI tools as too expensive or too shallow for your harder problems.
Sources: