Sharing AI progress in mathematics(openai.com) |
Sharing AI progress in mathematics(openai.com) |
Prediction: one of these is wrong and this (publicity stunt) will backfire.
Edit: don't tell me about lean. For lean to function as a proof certificate you need to represent the theorem correctly. Again: good luck doing that across such a broad swath of problems.
> Some of the unformalized results could have issues. We will endeavor to fix any such issues quickly. We are also exploring community-hosted repositories for these materials.
If they are all wrong, that's when it would backfire.
The question never if something works 100% of the time but how often it breaks and how that fits the need well. Solving one of these problems is a massive undertaking and accomplishment for the best minds, solving hundreds in a month but being wrong about 10% or something would likely not be the death knell you believe it to be.
With a deluge of results, having some human expert vouch that it even might be worthwhile would help. (e.g. see the link on HN yesterday, "Two Room-Temperature Antiferromagnetic Semiconductor Candidates" - I see it not worth looking at unless a subject expert vouches for it).
Assume the empty list you see is complete. :)
1: Author 2: Verifier
/s
It seems that OpenAI has a proof machine that keeps multiplying fruitful proofs!
I guess their job now is "Idea Man" and "Error Checker"? Kinda like (some/many) "programmers" these days.
>> may be thats what open ai did :)
That copium didn't last for what, three months?
"We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models."
To me, this is a take against progress so that mathematicians can keep their jobs. What would we do if, instead of math, we were talking about diseases? Are we going to keep diseases around so that doctors can keep their jobs too?
The ones with lean proofs could still be formulated incorrectly
Basically "Here you go, have fun with this, fuck all your demands, by the way we're gonna be releasing the model stay tuned!"
> At present, some frontier AI labs are testing advanced mathematical problems on proprietary models that remain inaccessible to the broader scientific community. Our recommendations are formulated with this practical context in mind. However, ideally, they would not do so. We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models.
To me, the issue is that the models are proprietary which are only accessible to a few people in 2 digits. It's not about progress but access.
It's been a while since I was reminded of this xkcd: https://xkcd.com/435/
How this maps back to math, idk.
> "I believe that AI can contribute positively in all of these directions [NB: exposition, community building, new directions of study]"
In SWE as well - this is what I do most of the day
They won’t because they don’t care and the only way it got this good is something along the lines of they trained on every mathematician’s codex sessions even if they opted out because they consider the thinking traces or output and metadata fair game.
These models are too expensive for broad access unfortunately.
Were in an unprecedented time where the value of knowledge is about to be crushed.
Have you ever heard of an S curve? Things will develop rapidly, then equalize. If they don't, we're at the singularity and I guess the end of time as we know it.
But I guess really bad things happen, cancer, radiation poisoning, torture, people have died in really horrendous ways, and I guess dying from some horrendous AI side effects is possible too. Yay.
The highest ranked would be:
| 22 | Hilbert’s tenth problem over ℚ |
| 29 | Unique Games |
| 31 | Anderson-model extended states |
| 37 | Spacetime Penrose inequality |
| 48 | Nonexistence of Landau–Siegel zeros |
| 52 | Baum–Connes |
| 78 | Abundance |
| 80 | Hadwiger |
| 87 | Bose–Einstein condensation |
| 92 | Two-dimensional entanglement area law |
At least put a disclaimer for the ad for this site, and maybe disclose how you came up with a total ordering for "top" open problems (vibes)?
> How problems are ranked. LLMs compare pairs of problems. A reliability-weighted model combines those judgments into the ranking, with calibration across model families. The model-family weights are OpenAI 1.00, Claude 1.00, GLM 0.95, and DeepSeek 0.90. These are modeling choices, not measured probabilities of correctness.
+----------------------------------------------------+------+---------+-----------------+
| Category | Full | Partial | Matched / total |
+----------------------------------------------------+------+---------+-----------------+
| Geometry and topology | 25 | 7 | 32 / 74 |
| Algebra, representation and category theory | 17 | 2 | 19 / 53 |
| Analysis and PDE | 11 | 6 | 17 / 40 |
| Number theory and arithmetic geometry | 4 | 13 | 17 / 117 |
| Probability, ergodic theory and dynamics | 11 | 5 | 16 / 37 |
| Combinatorics and discrete geometry | 7 | 2 | 9 / 34 |
| Theoretical computer science | 4 | 4 | 8 / 57 |
| Mathematical physics | 5 | 1 | 6 / 19 |
| Applied and computational mathematics | 2 | 2 | 4 / 8 |
| Quantum information and computation | 2 | 1 | 3 / 17 |
| Cryptography, coding, information and optimization | 1 | 1 | 2 / 26 |
| Logic, foundations and set theory | 1 | 1 | 2 / 18 |
+----------------------------------------------------+------+---------+-----------------+
| Total | 90 | 45 | 135 / 500 (27%) |
+----------------------------------------------------+------+---------+-----------------+> In a 2020 piece in the Notices of the AMS, I asked the following question: “If one human had an understanding of all of modern pure mathematics simultaneously, how much further would they immediately be able to see?” Six years later we are beginning to understand the answer to this question.
Discussed here:
To grieve, or not to grieve? - https://news.ycombinator.com/item?id=49919676 - Oct 2026 (156 comments)
I'm relieved that this time around, they have provided partial reasoning traces and prompts for a small number of problems. Do they now also share data with the other model providers..
https://github.com/openai/math/blob/main/preprints/Paired-st...
A Polynomial-Time Algorithm for Three-Machine Unit-Job Scheduling [1]
Since some people talk about small numbers that pop up in integer multiplication results, here a completely different number appears:
Theorem 1.1. Let an explicitly listed finite directed acyclic graph specify the precedence constraints on n >= 1 nonpreemptive unit-length jobs on three identical machines. There is a uniform deterministic algorithm that constructs a feasible schedule of minimum makespan. Given also an integer deadline 1 <= T <= n, it decides feasibility exactly and returns a schedule whenever the answer is affirmative. Both tasks can be performed in O((L + 2)^150020) steps on a deterministic multitape Turing machine, where L is the total binary input length.
That is some crazy exponent -- plus an interestingly old computational model to boot; not something that is natural to most of us. I have no capacity to check its correctness today, but I hope it is true purely for the exponent.
[1]: https://github.com/openai/math/blob/main/preprints/A-polynom...
[0] https://en.wikipedia.org/wiki/Unique_games_conjecture [1] https://github.com/openai/math/blob/main/preprints/The-Uniqu...
> Sure, mathematical history features a lot of incredible developments, like the invention of proof, zero, or the computer, and on the great problems our progress has been over timelines measured in decades or centuries. Obviously this technology didn’t appear today, but blurring our eyes a bit to combine the past ten years, with today a measurement of those developments, there is nothing comparable.
Cool! It's interesting to me how little my day to day work has to do with the big conjectures in math. Maybe I need to start posing some of my blockers as conjectures for some AI company to mine.
Look at one of their examples of an initial prompt: https://github.com/openai/math/blob/main/reasoning_traces/re...
Interesting that its only an excerpt. I wonder what else they include but didn't share.
https://wikimediafoundation.org/news/2026/10/05/openai-rogue...
No idea what it's so excited about, but it's cute that it "is." I for one welcome having access to a math buddy 24/7 that's way above my level but also always "willing" to talk at where I'm at.
Beauty can be appreciated even when it is vast, even when it is beyond one's comprehension. I don't think this release should be primarily viewed as an outcome of competition. Instead it is revealing truths about the universe that were always there and always beautiful, even if we hadn't seen them yet. I believe there are infinitely more such beautiful truths currently hidden and waiting for us to discover.
[1] https://www.nytimes.com/2026/10/06/science/openai-math-probl...
Like do you see the technology plateauing at the current level, do you expect progress will continue but only in mathematics, I'm interested to know why others are not concerned?
109. Integer multiplication below n log n
Surprising that this is possible.
158. The Euclidean plane cannot be colored with five colors.
Only 6 and 7 remain!
376. Universal computation in forced Navier–Stokes flows.
Morning coffee proven turing complete
LMAO, I don't think I ever saw such a small number in a CS result.
Like there is somehow redundancy in a fourier transform that makes it sub Linearithmic?
Which low and behold ->
130. Fourier transforms below n log n.
Very surprising result though! Multiplication is easier than sorting.
Does anyone have an intuition to what causes it? What happens at these large scale (or very small)?
What I mean is progress in math comes from having gained a deeper understanding of the problem for subsequent attack of more problems and IMPORTANTLY applications! Right now the first one is trivially satisfied (given oai maintains some memory across models) but the second one is not! It's generating proofs faster than anyone can validate and so only the model can use these. Consequently if it just keeps doing more theory it's... not very helpful or at least not optimally helpful. This is just bragging rights for now.
More "application-oriented" research would be awesome, where it tries to achieve some desirable effect and then produces relevant theory and experiments around it. Fields like CS, Physics, Chemistry, etc. This would also benefit a wider section of the population rather than the 10 people who understand most of these proofs.
It's fine if not, but it'd be great if even just one of these helped us solve a long-running problem.
I'm not sure it's clear right now.
No clue why I'm downvoted for this, HN struggles with truth-seeking on these topics.
We cannot have them rushing to publish amidst tons of confusion, rumors of threats/scooping and outright plagiarism of existing work (by failing to cite said work).
If they're going to participate as scientists in these more rigorous fields, they're going to have to match that level of rigor, not lower it to the disastrous low that ML research publication is at.
There's no gatekeeping here!
Said communities are doing that all on their own by caring about what AI is doing rather than just focusing on their own thing as they did before AI. It's a serious kind of envy IMO.
How do you know? Seems statistically unlikely with 720 problems, most of them well known
Seems like a lot of PHD students are doing to have to pivot the entire structure of their PHD studies? Or just produce something which is already written by OpenAI?
Can we get a number in Blackwell GPU-hours, kWh, or some other compute-scaled metric?
That doesn't sound right
But having so many of them at once? Damn. We really live in the future.
> A Unique Games instance has a finite vertex set, a finite alphabet K, and a nonempty list of oriented constraints e = (u_e,v_e,π_e), where π_e is a permutation of K. A labeling a satisfies e when a(v_e) = π_e(a(u_e)).
I'm sorry, what? I admit it's been quite a few years since I've thought about the Unique Games Conjecture, and I never dug that deeply, but this part is very, very elementary graph theory and notation. So let's unpack it.
1. e is maybe a name of a list.
2. The elements of that list are tuples, where each tuple is (a vertex, a vertex, a permutation). So e indexes into the list and u_e is the source vertex for the e-th constraint in the list called e. Thanks.
3. a is a labeling. I'm fairly confident that, by "a labeling", they mean that e is a function from vertices to colors, where the colors are the elements of k.
4. That vertex coloring a satisfies the list e, when, for, um, an index e into e, a(v_e) = π_e(a(u_e)). But this isn't for all e, it's for some e, and the goal is to count them.
So maybe e isn't a list? Maybe e is a constraint that is represented as a tuple, so e = (u_e,v_e,π_e) and u, v, and π aren't sequences at all but are, in fact, the trivial unpacking functions that unpack the pieces of the tuple.
Reading this stuff is pointlessly painful, and it's extremely easy to make mistakes when being sloppy like this.
If this were my paper, or if I were trying to train a model to write math, I'd want something like:
A Unique Games instance has a finite vertex set V, a finite edge set E = (V × V), a finite alphabet K of possible vertex colors, and a nonempty list of oriented constraints. Let Π be the set of permutations of V. Each constraint e is a tuple in E × E × Π, where we write u_e ∈ E for the first element, v_e ∈ E for the second element and π_e ∈ Π for the third.
A vertex coloring a : V → K satisfies e when a(v_e) = π_e(a(u_e)).
He spent years formalizing his sphere packing theorem because the proof (human produced) was already beyond the ability of peer reviews. Now his formalization effort likely can be easily reproduced by a model. However one should read his experience about what a formal proof is: often the problem is the statement not the proof. The example he gave is the Jordan curve theorem. It's actually quite challenging to formalize the concept of a planar curve (there are space filling curves). So it is not necessary that someone can look at a formal statement and say aha it is about a planar curve, unlike FLT where there is not much problem in recognizing what the statement is about.
We are not far away from the moment where these models will be restricted, and sharing the results will be done more carefully.
Who is "we" here exactly?
Not really? We are at a point if an AI today can solve it, it can be stepping stone of understanding something deeper to tomorrows AI and it continues. Sort of like our limitations doesn't matter. Obviously there are many scenarios in this recursive loop but saying it isn't much progress is not how I view this as
Look at that, taxpayer funding was cut and a private sector solution came in just the nick of time, far accelerating the holding patterns we’ve been in for decades
Humanity doesn’t need all iterations towards the blueprints, the blueprint is good enough, we all stand on the shoulders of giants
Proven math theorems are tautologies.
Before I cope, I’ll note that there are plenty of “doom” scenarios that do not require any improvement in capabilities from what we had before this latest unreleased model. We’re at the point where a determined bad actor with enough compute could compromise critical infrastructure in a way that results in casualties, where this actor would not have been capable of such without LLMs. This may not sound like Skynet, but I don’t see why it makes a difference if I’m one of the casualties.
With that in mind, here is the cope: first, mathematics is an inherently verifiable domain. An LLM can use tools to determine with absolute certainty whether it is correct, and an independent third-party could review and confirm. All of this can be done without any interaction with the physical world or with other minds.
Second, OpenAI is able to marshal compute at a scale that an individual mathematician can only dream of. It’s possible that these problems were lower-hanging fruit (in relative terms), such that they could be resolved simply by throwing a ton of compute at the problem guided by an intelligence that is not itself remarkable in comparison to a human.
Third, none of these problems are solved in a vacuum - the reason OpenAI chose these problems is that they are widely discussed and many people are working on them. It’s possible that someone else was close, and OpenAI only contributed the finishing touches. (This wouldn’t need to be plagiarism, to be clear - people publish their work!)
See:
> The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking. (TFA)
In looking at this over the past hour, I haven't seen clear evidence one way or the other. Some of the stuff is highly unexpected (like the multiplication algorithm), but counterexample-y, and about the rest the professional mathematicians online seem to have a consensus that it's not "breaking through fundamental obstacles". I suspect neither of us is competent to judge that.
Superhuman tenacity is not enough on its own to pose an existential threat. If it showed the same capacity for judgment, inventiveness, and decision making in the messy problem space of the physical world, I would be more alarmed. There have been experiments where an AI is given control of managing something like a vending machine and it always ends up a mess. AI has come a long way, but certain problems seem as difficult as ever.
When AI becomes more capable of navigating practical problems without human intervention, I will start to be concerned. Enslaving humanity will involve taking a lot of calculated risks that tenacity alone cannot solve.
Given the past rate of progress, why not start being concerned now? It's a bit like the economist saying that the optimal number of flights to miss is not zero. If you keep landing short on your estimations for how far this technology will go, next time you should err on the other side.
And regarding Vending-Bench 2 (https://andonlabs.com/evals/vending-bench-2) my understanding is that models do pretty well on it now.
[0] https://dank.systems/posts/2026-09-15-ai-bear.html
[1] https://gowers.wordpress.com/2026/08/12/what-sort-of-maths-a...
We'll have plenty of time for this, while living off UBI.
But they’re already extending into politics, military, journalism, art, and many other fields that aren’t verifiable in any meaningful sense of the word.
Is there any specific cognitive task that you are willing to bet that AIs won't be able to accomplish in the next 5 years? Because if not, I'm not sure we're disagreeing about predictions.
Or is it simply that you feel bad for Mathematicians.
I really think the only place people disagree is that they don't actually think it's possible, they see it as hype or doomerism. I can't find any good reasons to rule out that the companies could actually achieve what they are trying to so I think they should be stopped.
If doom is ending up with grey goo / paperclip maximizers, or SkyNet, then I don't think doing mathematics is evidence of that direction. Partly because LLMs are quite apparently dumb in many ways, and for math specifically, they need a formal verifier (Lean) which "gamifies" math.
If you mean bioweapons or cyberwarfare, there's nonzero risk, but not orders of magnitude worse than other global risks. Climate change, nuclear weapons, monoculture food, etc.
I'm far more concerned about overall trends in AI development and usage. It's accelerating wealth inequality, social isolation, attention capture, surveillance states. If we end up in the Matrix except the admins are humans and the simulation is hyperoptimized TikTok, is that AI-driven doom, or is it just an inevitable outcome of modern tech?
So in other words, since deep learning is algorithmic research, we are now in the RSI era.
"Surprising" is a, well, surprisingly high bar to clear, and requires thorough understanding of not only the paper, but existing work in the area. ("Novel" is tautological.)
How did you determine this in 1 hour? Are you a researcher in multiple of these areas?
Can you give an example, or explain more how you came to this conclusion?
a) destroy chess and make it a pointless endeavour,
or
b) make humans much better at chess.
I do think it also took some of the magic away from chess, and Lee Sedol has said something similar about go.
So did it destroy it? No. And maybe you could make the argument that it got more exciting in some ways, but I think it sort of degenerated into a spectacle and it's just not as interesting as it used to be, and I think computers have played a role in that.
The issue I see is the list of abilities that AI can't do is shrinking at a rapid pace, and its capabilities are growing at the same pace.
So basically I think that the future is getting pretty weird because we are building really powerful tools, though these tools are precisely what allows us to prosper in that future.
I have a hard time taking statements from OpenAI about their own product, that they are trying to sell to people and make money, seriously. I take these statements as they are greatly exaggerated or even straight up lies and propaganda.
That said, I think results like these are mostly annoying more then anything. They spent a lot of money, used up gigartiuan amount of compute, to ruin a puzzle that mathematicians were tackling. I am mostly unsurprised that if you spend a trillion times the energy that a team of mathematicians would, that you get maybe 1.5 times the results. I see a future where that 1.5 times the results may go to 3x but not much more. And if that, then I will be more annoyed.
Yes, human beings can do more now in some ways but...to put it poetically, I think there will be no more heroes like Einstein and Newton of the past. Now it will just be someone cleverly turning the crank.
Yes, we still admire Usain Bolt even though we have cars...but maybe the admiration is a lot more trivial than if we did not have them....
Personally, I think AI is a grand mistake.
Why do you think the world to date hasn't been taken over by evil genius mathematicians? Can you extrapolate from your understanding of the answer to that question?
There will be a lot of job loss unquestionably, in the same way that automation reduced manufacturing jobs and farm payrolls.
At the same time we have to put what AI can do in perspective.
Intelligence is a broad grouping that includes concepts such as knowledge, skill, experience, and wisdom.
AI has incredible knowledge and in many areas approximates experience and wisdom.
But wisdom is harder to formalize than knowledge and skill.
For example certifying a college education relies mostly on the ease with which we can verify/test knowledge.
To some extent advanced degrees try to certify maybe wisdom and experience.
In my very personal opinion, wisdom and life experience should give humans an edge for a while to come.
Additionally, I do feel that the more an individual lacks better than average wisdom and experience, the harder it will be for that person to compete with AI.
Also, on the bright side, the average human will continue to prefer to interact with a fellow human in many spheres. That will also act as an upper bound on AI and robots taking every job.
Either way, I do think this transition will be painful. I don't feel it has to be apocalyptic.
But the world has been an especially volatile place over the last 10 years.
So when you add that existing volatility, to the upheaval from the AI transition, it would not surprise me if the transition results in violence.
But, without the pre-existing volatility, and if humans were capable of generosity and love at scale, I see no reason AI can not be absorbed into society with net gain.
I guess to summarize, I feel this tech should be a net gain and to the extent that it isn't, it will be because of flaws deep inside of humanity itself, not because it had to end in chaos.
In other words, I feel fear, greed, anxiety, and competition -- all our base instincts coming from all sides, will be what determine the end result of AI moreso than AI taking everyone's job.
It has to feel awful to be in this position.
Very hard question.
Your work makes you one of the very few people who really understands the problem and solution and its significance.
I guess the only answer is to adapt with the tools. If we can't do that, then yeah, we're in trouble.
Precisely what all NLP researchers and the ML community at large did in the last few years: embrace the frontier and realize that attention is all you need.
I did get my PhD...before AI. And my honest advice (to myself back then, even) would be: quit the PhD, become an electrician, and work hard to buy a tiny house in the middle of nowhere to watch the world burn in this madness.
with non-polynomial side being represented as the frontend programmer's constant need for more performance to do the same task...
I’d posit that more people are playing more and learning chess than ever before, thanks to networking and AI assistance. And computers have only beaten us at computer chess. Human chess is always an experience for learning about the other person, or flipping the board and walking off in a huff.
I don’t know so much about Go and it’s not surprising that Lee Sedol became pretty demoralised, but the generation coming after him alongside computers are going to see new possibilities that had gone unnoticed in purely-human Go, extending the game for everyone.
As in..to be dominant? Why would an AI try to dominate? What would give it purpose, or is this a purpose via misalignment scenario?
My approach would require custom engineering for every different sequence we'd want to target. With CRISPR, you just "program" the system with a guide sequence, you don't need to do massive engineering to solve a protein design problem.
:)
In this case, the tenure is gone, and OpenAI has increased their valuation
Right. Before all the AI disruption, pure Math traditionally welcomed anyone who wanted to study its esoteric proofs, right? I remember all the excitement of the average Math enthusiast casually reading Wiles' proof over coffee.
Bottom line is, the relevant people can still understand the generated proofs. The disorienting part is they are a little slower than they'd like, but they'll get there.
That is 1. immediately technically possible, and 2. realistic.
If you need a source for 2 I'd suggest you open any history book.
i get a lot of skepticism on HN by the same crowd that has been wrong about this tech for about 4+ years straight
Why would a biolab capable of making something like be unregulated? And if it definitely would, isn't the problem with the biolab?
It feels like all these scenarios are leaving some gaping holes in our security infrastructure that have nothing to do with AI.
Mathematicians aren't exactly known for being well rounded.
Yes, you can be an amateur mathematician who manages to avoid such classes and may not know those results or objects, but if you haven't read and written down the name enough to avoid habitually misspelling it, you are outing yourself as a meat proxy unless you are dyslexic.
There’s about a 1 in 10 chance when someone spells it (or says it) they say Herbert.
Even in situations where they just read it or I just said it.
I’ve had Herbert soccer trophies, health insurance cards, etc.
The mind fills in a lot of blanks and doesnt always get them right.