They don’t have any more patience for this.
It looks like people are enjoying themselves, having fun with the new results, and generally doing all of the things you say "science" is supposed to be about. So what's the problem?
https://www.youtube.com/watch?v=LKiBlGDfRU8 https://www.youtube.com/watch?v=shFUDPqVmTg
As the author of the post points out, there is no way this is “the best they could do”. It’s a write up that didn’t involve someone with the math + communication skills required to clearly explain the result.
Nice idea.
(See https://agmai.org/general-sep29/ for the recommendation in question.)
People seem to have very misguided ideas about why OpenAI is doing this at all. It is not to brag or to torture mathematicians. It is an eval. OpenAI is known to be willing to pay large amount of money to get a good eval, think FrontierMath. FrontierMath is now saturated, so they need a replacement eval for math. Open math problems are actually a fairly good eval, although a proper eval is better (eg FrontierMath has known difficulty and have tiers from 1 to 4).
Mathematicians would prefer if OpenAI didn't use open math problems as an eval, but OpenAI is not obliged. I actually think OpenAI wouldn't point AI to open math problems if unsaturated FrontierMath Super Duper is available, as it just angers mathematicians, but such eval is not in fact available. Given OpenAI used open math problems as an eval, they could just throw out the result (this is in fact better as an eval since it will keep problems useful longer), but mathematicians preferred to see the result. So OpenAI released them.
People who are invested in the idea that we've invented a general intelligence, now, which includes all these companies that are literally financially invested in this claim they are making, will tend to believe that its results can already be trusted in domains like this. Some mathematicians seem to believe some of the proofs written by their models, and some, like this one, don't. I do think it's valid for an expert to push back against the claim that the best use of their time right now is to verify the poorly written work of everyone who's claimed to solve the problem
If you just strip mine the answers and Sam Altmans magic button solves 100/100 problems, what's next? Who is left to come up with a new interesting question for the magic button to solve?
Lastly, life and the present moment is all there is, if there is no enjoyment in anything we do, then what's the point of all the "living for ever" Altman et al want to achieve.
We will live forever to read boring papers generated by LLMs? Literally sounds like an eternal hell.
In fact, I can't remember a time when I was more excited about the future of science. This could herald an end to the replication crisis, and kill off bullshit science completely. The danger of course is that we end up with two companies effectively dominating cutting edge research in every field, but it remains to be seen if that's even possible given the pace of improvement in open weight models.
He's tracking the community progress on sub-n log n multiplication. OpenAI started with 1 - 1.63e-55. The result has been now improved on 115 times, and the current record is "rohanarun"'s 1 - 9.87e-5. I'm sure by tomorrow it'll have improved again.
Does this look like people aren't having fun? Does it look like they aren't discovering stuff? It looks like it's spurred a cascade of interesting community activity. It doesn't really seem much different from what happened with the twin primes conjecture. Isn't that supposed to be the point of all this?
This is like "no one is forcing software engineers to use AI tooling" or "no one is forcing you to show your ID in the airport" or "no one is forcing you to own a car in your small midwestern city" - there can be no law requiring something and the practical consequences of not doing so can be so painful that you're effectively forced anyway.
As someone who uses LLM tech occasionally, this is why I prefer using open local models. If I’m making myself obsolete, at least I’m not making some asshole richer and their closed model better.
I just imagined that instead of math papers, they released 700+ feature length films, and the only way to tell if one of them is any good is to watch it in its entirety.
That feels pretty unappealing to me.
I know it's the same for human made films, so what's the difference right? But those are good enough most of the time that it's a decent bet, and the people that made them had real skin in the game.
Contrast that with something made by a nondeterministic slop machine with no skin in the game where small details can be off in a way that's jarring. Right out the gate I have an aversion to committing that much time to something that very well may waste it.
However, there are always smarter, hungrier people out there and this is a buffet.
Some output is going to be wrong or incomplete. I am willing to bet even those have nuggets that can be used elsewhere.
Like people enjoy racing in front of a stopped train? As soon as they turn on the engine again, they will run you over. The questions that remain will be only the low value ones, not worth the effort to vacuum up.
So no, the smarter, hungrier people are not the ones that are going to swoop in. It will be the most desperate.
> Some output is going to be wrong or incomplete
This is a very human take on the situation. No, the Lean proof is not going to be wrong, and it will be incomplete only in the sense that OpenAI didn’t try to push the results further.
There's a "Silicon Valley-ism" for you. We offer a thing in whatever form we want and people "who are passionate" will gobble it up, should gobble it up, 'cause they're "passionate".
Yes, the situation sucks overall and mathematics as a whole is in a turbulent time now.
But it also sucks when mathematicians, who are considered experts on a particular problem, refuse to engage with breakthrough results about that problem. #2 above is still true regardless of where it came from or how hard it can be to absorb.
While reading "The Mathocalypse" post [0] by Scott Aaronson, Scott described his wife Dana's reaction to one of the newly solved results in her primary domain of expertise, on which she'd been working for decades.
After her initial shock, and annoyance with the format/style, she decided to start using Astra - for the first time - to help her understand the new result. And he reported in the comments that she had made a lot of progress understanding it in one day, and may be even excited to give a talk about it!
That seems like a much healthier attitude towards these new results.
Yes, everything else sucks about this messy period. But there are still diamonds (in the rough) in this drop that perhaps should be looked into. If the author is too busy, perhaps one of their students can take a look? Someone will, eventually.
It might be a shock for you but they are very few in numbers. Most of researchers I know are always busy with something. They cannot just drop other responsibilities for something like this. They will take their own time getting through the proofs (if they want to).
> That seems like a much healthier attitude towards these new results.
Another thing to consider is not all mathematicians are from US or with good funding. The PI or graduate students cannot afford to pay 200/month.
TFA was about a niche topic that OpenAI doesn't have in-house expertise in.
Otoh Aaronson is the co-author on Lijie Chen's (reasoning lead at OAI) top cited paper. OAI have deployed their resources more effectively against UGC that some of their staff are already familiar with
I wish they take a bit of more time to communicate the findings effectively.
I don't believe this myself. But I do believe that if you've formed your very ideas about what is good and desirable on the basis of a culture that has held certain values dear for hundreds of years, and have fought against every doubt and difficulty in life for decades to mold yourself into that image, that it does not 'suck' that you are unable to adapt to a new reality overnight.
Very few people that would love to be craftsmen would love to be factory foremen. It is far too insensitive to the human experience to expect people to just deal.
If some mathematicians complain about OpenAI sucking, that is fine actually, and if others are more “mature” about it, that that is fine too. Neither of these reactions should be at put as an equivalence to the blame OpenAI deserves for this stunt.
If you did’t bother to write it, I shouldn’t be bothered to read it.
Perhaps AI agents can have their own publications and magazines where they are the chairs and associate editors and reviewers.
If AI can solve such grand, outstanding math problems, and mathematicians argue these pure math problems are important, what’s the problem with them needing to read the output if they want to understand it?
Academics have always been required to engage with hacks and cranks to some extent; the deluge of AI proof writing has only exacerbated the problem.
English speakers generally use this word in a very broad sense “and now Netflix is forcing ads on paying users”, “because there was no sink, I was forced to drink the whole thing”. It is only when you are literally describing a crime where this word has this strict meaning you are alluding to.
This part I don't understand. Not that anyone should read the entire Lean code of any proof, but if the statement of the theorem to be proven in lean seems to be correct, then I would think there would be at least some interest if in fact there was a formal proof (which might or might not correspond to the written proof) of something I was working on. That to me would be interesting. Or you are saying you doubt the validity of the formal proof, which would also be interesting. But saying it is of no consequence doesn't make any sense to me.
It would be like trying to look at a completed video game's assembly code, being told that it was call of duty, and then being asked questions about the high level code architecture.
AI models are perhaps unsurprisingly good at low level translation (see the progress being made for decomp games)
These models have surpassed human capabilities at math/machine code, but they can't "simplify" yet - in part because they don't have the same need to due to their comparative lack of cognitive constraints. AI Slop code is getting better, but it takes time. At the moment, its embarrassing frankly. It will come eventually, but right now OpenAI is not handling this with the care, respect, or concern that it deserves.
Work this abstract almost certainly has no value outside the community that is (was) interested in the result. OpenAI should engage with the community to realize the value (beyond PR).
I sympathize with those who've worked on some problem for years and now don't have something to work on; it's been a part of their identity. I also especially sympathize with those whose career tracks and plans were thrown in disarray.
That being said, I absolutely cannot understand how one can't be excited and happy and enthused about these advances in one's field. Assuming just that the ones with formal lean proofs are actually true, these are reportedly huge advances. Even if folks don't understand it YET.
A few papers have been retracted, but it looks like many are withstanding intense scrutiny. Lean is making the results more likely to be correct, but I think making them harder to understand.
The world has changed and you’ll know a math department is making a serious attempt to adapt when it teaches a required Lean course in freshman year.
I know the model that produced these proofs is still private, but it’s worth a shot tackling the proofs with the current consumer-available frontier.
...
>So, no, I will not be sending Sam Altman a bottle of whisky anytime soon, nor I am planning on spending my time reading through that paper and trying to make sense of it.
Think about a hypothetical circumstance where we get radio communication with some aliens on another planet. They send over tons of math to help us advance our tech, we know the math they're sending us is correct, but their explanations are really hard to work through because they aren't humans and the math is so different from anything we've done. Should we whine about the results they sent to us and refuse to engage with it?
The major discussion is about whether this will in fact advance our tech.
In the Three Body Problem, the alien race shows us “miracles” in an attempt to discourage us from the pursuit of science. Consider that possibility.
it won't be long before there's no more low hanging fruit like this to complain about, and the writing / explanations of the results are superhuman as well
separately, i really liked the author's denial-of-service analogy. super useful practical framing
For what's worth it, this batch of results are not very tight intentionally by OpenAI and promising mathematicians already started to consume them and improve the results, while someone is still complaining on it.
The lack of emotional maturity and empathy is very on par with my experience thus far.
> This is why I generally avoid using AI for mathematics (I am happy to ask LLMs to consolidate information for me, or to generate a useful infographic, or to proof read an email, etc.)
In other words, the author is OK with using LLMs to replace data analysts (that could consolidate information), to replace graphic designers (that could generate infographics), and to replace editors (that could proofread an email). But don't you dare use LLMs in their mathematics.
Excellent point.
I don't blame the AI for this - I blame OpenAI.
Literal slop grenade (see https://fortune.com/2026/09/17/shopify-tobias-lutke-ai-slop-...)
Imagine being someone who is working on one of these problems. You have no good guarantee that the problem was solved, but you will have the horrible homework of reading the AI slop. Also, if you do have something interesting to say about the problem, people will have less enthusiasm about it now
When I am tasked with reviewing that code. First, I don't know if it's valid. The person who wrote it doesn't know if it's valid. In order to validate it I must step back and understand the full problem space. Then, when I ask for revisions or clarifications, it's seen as either
A - Slowing progress, being resistant to change... or... B - Thanks for catching that (Claude fix PR 532 with the review comments)
It's 100% removed the enthusiasm.
Perhaps, there is a point where we just "give up" the understanding and accept AI output as the ground truth, because the sheer amount of generation is too much for our puny human minds to comprehend, and a lot of the times it IS right, even if a little wonky.
It’s clear from the author’s tone about lean, emails, infographics, etc. that he thinks automating those away is fine. Why should math be any different?
In fact, the academic system is a kind of worldview created by humans. And as it is shared and the community grows, the problem will gradually become more complex. Because when a discipline develops sufficiently, just as in a mine where rich veins are easy to extract early on but become very hard to extract once much has been dug out... in that sense, as things gradually become more complex, once a certain threshold is reached, won't scholarship surpass the limits of human understanding? Of course, scholarship is entirely for humans, but at some point the system itself may face its limits, and then wouldn't it again reduce the existing normalized minimum within that discipline and establish a new normalization of a new logical system?
In my view, perhaps for very complex work like today, AI will do it, and then there will be work that normalizes and further simplifies the results of that AI. Then, coming back to the human fold, if humans create the initial skeleton, the LLM will learn that again and it will become complex work again, and won't this create a continuing cycle?
I think verification and understanding can be separated. If the proof targets a correctly formalized proposition and passes a reliable proof checker, isn't it valuable? We have obtained knowledge justified as true, but there is simply no new theory that understands that knowledge. As was the case with the Four Color Theorem...
I am always curious what shape the newly compressed new discipline will take. At that time, I hope even people like me, who are intellectually behind, will be able to learn that discipline.
No.
But its worth a lot less than one that can be understood.
They really should be trying to partner with the mathematics community to add maximum value.
Their current approach is reckless and risks doing more harm than good.
Sorry, but I don't get why mathematicians are so upset. Like, just accept the knowledge and insights and acceleration in your field! If it isn't "fit for human consumption" because an AI produced, okay... it soon will be explained ELI5 by even better models.
The same could go for mathematics or any field; there are lots of people who enjoy the process and aren't satisfied by being handed and opaque final result
Their field is at the stage where the humans are “debugging” the AI slop.
They are still imagining how to escape from having to read the generated code. Hopefully they find a way.
Most software developers (or people in any field, working for anyone) rarely own anything they do at work.
And the entrepreneurs running their own companies (which there's an explosion of atm largely because of AI) do indeed "own" the higher-level products and things they're producing, even if AI writes the code.
What is supposed to have changed?
This is a strawman. OpenAI didn't say they are expecting all mathematicians to read the solutions, incomprehensible or not.
So mathematicians are upset with OpenAI for solving "their" math problems. Software engineers are even more affected by AI, yet mathematicians seem to be reacting more strongly. I don't get why.
It looks like the spent $20 on writing the actual papers.
If they actually wanted to do good for the world, they wouldn't have released these as the slop grenades they are.
In their current state, they are actively damaging the mathematics community.
It shows a lack of respect and care for the impact that their technology has.
It shows that they cannot be trusted for things like private data, AI safety, and company partnerships.
In math/science, repeatability and review are critical to the process.
The right way to handle this would have been to work with the mathematics community to co-develop and create meaningful proofs rather than slop grenades.
If they proceed in the current state, we'll just get a bunch of spaghetti math that won't do anything for helping people build an understanding.
Maybe some day, we won't need people to understand things, but that's certainly not the case at the moment, and likely won't be for several more years.
Perhaps this is projection, and the staff at OpenAI doesn't understand their work anymore? Not a great sign regardless.
The mathematics community can finish the job OpenAI started. Or are you saying the community has no incentive to do that because there is no reward/recognition for doing that?
They didn't just "start" it though - they released slop papers.
That's finishing it - not starting it as far as scientific publishing is concerned.
> Or are you saying the community has no incentive to do that because there is no reward/recognition for doing that?
Not exactly - but that is part of it.
I think what would have went over better is:
1. Immediately announce a solution has been found.
2. Do not publish the solution.
3. Put out an open request for anyone with experience in the area who wants to get involved to help collaborate on a construction and human-comprehensible paper. Accept anyone who can demonstrate potentially useful work/experience in the field/problem. Share the solution with them after they sign some kind of NDA that they won't independently publish or share the solution/work.
4. Work with people until a paper is ready (I mean actually ready - not the kind of slop that they released).
5. Publish. Include names of everyone who made meaningful contributions to the paper (not just the proof).
EDIT: Notice the incentive with my proposed second path is that it gives OpenAI an incentive to improve the interpretability of its proofs. This is a good thing! The maths community would be thrilled to actually gain understanding from such releases, and OpenAI would be happy because they could more quickly and independently publish their results. At the moment the "value" of their mathematics research "product" is low because of the lack of this interpretability, and this current approach is simultaneously destroying the opportunity value of the community as well as the incentive for OpenAI to ever improve on what's missing.