If a subset of mathematicians, use AI to condense timelines focusing on goal 6.2 exclusively and make rapid progress and reach a proverbial inflection point — one where value proposition of the using this new normal is too enticing to give up — everyone will ask: "This thing is so awesome. Why should I care about your values?"
> My own suggested rule of thumb: if the authors cannot convincingly demonstrate that they are able to give a clear, expert-level talk on their results, one that is correct and properly attributed, then the result should not be published. A proof that no human can properly explain should be viewed as incomplete, even if it has been formally verified.
I think these are non-trivial epistemology and science theory problems.
In that setting, field experts working at the bleeding edge are so advanced that non-experts literally can't understand what they're saying at all. So there's a whole class of specialists, "synthesists", that specialize in gaining approximate understanding of the experts' work for the purpose of communicating it to outsiders—perhaps wrongly, according to the expert at least, but hopefully more productively vs the unmediated version.
I'm deeply suspicious. I do not yet have a concise statement for why, but a lot of literature on the sociology of knowledge work sort of points at my thoughts.
Section 5 of the Thurston article cited by Tao touches the elephant. Raduchel's article on the economics of software [2] also touches it.
I've tried to put words to this for a few years. I think I'm just going to start writing versions of it as see if that helps me shape the thought into something more concise.
So, in the spirit of this article's style, here are some postulates:
1. There is a sociological process happening in the production function during knowledge work.
2. That production function and the associated sociological process spans years or even decades, and must outlast many of the artifacts that are produced during the early years of the function.
3. You cannot get the right lines of code or the right theorems proved without running that sociological process alongside the artifact production process.
4. It is impossible to completely separate the sociological process from the artifact construction process. If you just iterate on artifacts then too much of the required hidden state is lost to make progress in the right direction. This is true even if you include distilled artifacts capturing pieces of the sociological process (eg meeting notes, documentation, commit logs, prompts).
5. So you need that sociological process, or something like it, to still happen.
6. For a lot of knowledge work that process plays out in extremely high-fidelity social interactions [3] that we have not yet captured in the datasets that would be required to reproduce those dynamics.
7. And even if we do collect that data, our current architectures and training algorithms and hardware would be useless given the size of the datasets.
So: the technology today gives us the ability to iterate on the production of artifacts. But it does not sufficiently simulate the social process which gives rise to the Right artifacts.
This isn't exactly what I actually think, but it's a version of the thing that I intuit when I watch heavy use of AI in both software projects and formalization projects. And simulating that process feels way harder than people are currently assuming.
[1] https://arxiv.org/pdf/math/9404236 Section 5.
[2] https://www.nationalacademies.org/read/11587/chapter/11 pp 166-168.
[3] there is a reason we still gather in-person around white boards, and why doing so is more crucial for some types of work than others.
Properly explain is an enormous grey area. Soon, I think, there will be proofs of results that are verified in Lean that are so long that no one will be able to “properly explain”. I don’t think they should be discarded.
Resolution of singularities is a famous theorem of Hironaka. Abhyankar claimed that no one truly understood the proof of the theorem. He said that he and Zariski couldn’t get through the paper with a full understanding. But everyone accepts this theorem as being correct.
Just burn lots of tokens on the frontier model of your choice to let the AI find a high-level argument why the four color theorem holds. :-)
--
Seriously: since there exist quite a lot of readers on HN who are both hardcore into AI and mathematical problems: This is a challenge for you.
I am looking forward to seeing an announcement of a novel high-level argument why the four color theorem holds on the first page of HN in at most a month. :-D
Hmm, doesn't it take an expert to explain why those cases are exhaustive, and why the code that checked them is correct?
Tangentially, I'm not a mathematician but I wonder if one "opaque" proof that is too complicated for anyone to understand, but that we know is correct via formal verification, might end up being built on with "transparent" human-understandable proofs. For example, it's my understanding that there are many conjectures that have been proven true conditional on the riemann hypothesis being true. In that case, an opaque proof of the riemann hypothesis would enable those conjectures to be known and built upon
To your first point. There a large number of cases that maps can be reduced to. Very few people have checked these reductions themselves. In 50 years there will be no human alive that will have checked the reductions by hand. Do we then discard the theorem? More importantly, do we trust the people that claim to have checked all the reductions? There are hundreds of cases. I trust a computer verification much more than I’d trust human verification. Humans will likely make mistakes due to the tedium. And some will claim understanding of all cases but be wrong in their understanding in some of the cases.
But the point is that pre-AI it was already the case that famous results were published that very few could understand or digest. I think it is reasonable to expect that we will soon be at a point that Lean says a theorem is correct but no human can or will ever understand the proof.
What if Lean verifies Mochizuki’s proof of the ABC conjecture. Do we disregard it becuase no other mathematician understands the proof?
It would work better as a bar for hiring, rather than as a bar for publishing.
We'll end up with incomprehensible math because comprehensibility isn't rewarded. No one is going to get a Fields Medal, or tenure, for digesting someone else's results.
The incentive will be to be able to publish in a top tier journal. I suspect what Tao is advocating for is having journals reject such manuscripts.
> No one is going to get a Fields Medal, or tenure, for digesting someone else's results.
I'm sure no one gets a Field's Medal if others can't digest their results.
He says it shouldn't be able to published if they can't explain it. Publishing it is the reward.
Science,at its core, does not care about the credentials or institutions. It cares about the results and to what extend they can be falsified.
This feel a bit like "we know all about physics, we can only get more precise" - moment
I noticed a long time ago, that the more people focus on trivialities like typos when arguing against someone online, the more compelling the original argument is. Basically, bikeshedding.
The most compelling evidence of the compelling nature of the original argument is when the most-upvoted reply is a joke or a meme. That's when you really know that those responding have nothing else to say. It's a white flag being run up, or the dog turning over and exposing its belly.
Some academic cultures have a tradition of formal debates. They are based on the premise that an educated person should be able to argue convincingly for or against any idea, regardless of whether they believe in it. A natural corollary is that you should not let convincing arguments convince you, as the merits of the argument have little to do with the merits of the idea itself.
LLMs have made the situation worse. People's ability to generate convincing arguments now greatly exceeds their ability to evaluate the value of ideas.
If Amazon uses AI math to come up with better routing, the cats can benefit from cheaper delivery fees just as much as humans can. No understanding needed.
The human brain is being obsoleted, soon thinking is going to be a recreational activity like weightlifting. If you want to think as a hobby, that's fine, but most people will be free of that toil of unwanted brain labor.
https://arxiv.org/abs/math/9404236
He wrote it in 1994.
He writes about how he almost "destroyed" a subdiscipline in mathematics by becoming so good at it that he outclassed everyone. PhD students were advised to stay away from the whole field.
When he discovered this, he realized his error was that he was focusing on producing results, and not focusing on explaining his thought process. It's that thought process that is valuable in advancing the frontier - results alone won't do it. It didn't matter how many theorems he proved, if he was the only one who had the mental framework in mind on how to think about the whole field.
I'm sure we've come across abstruse books where every theorem has a rabbit being pulled out of a hat, whereas other readers find it intuitive. It's because the latter has developed a mental model for the discipline, and you haven't.
So he set about slowing down, and focusing on holding lots of seminars where he worked with other mathematicians to explain the thought process. Eventually others started publishing proofs of key theorems.
When people publish in a journal, they are not merely doing it to show the result. They are having a conversation with other mathematicians. If they cannot explain their own proof, they're not having a conversation.
This is why even decades after the Four Color Theorem was proved, plenty of mathematicians don't consider it "mathematics".
Useful thought, rather than hobbyist thought, seems destined to be the exclusive domain of silicon.
I don't follow - are you surprised that mathematicians have social rules on how they interact with others?
You're definitely welcome to set up a journal that takes whatever types of papers you deem acceptable. It's not like they're preventing the dissemination of information by taking this stance.
Personally, I wouldn't hire a SW engineer who only showcases output from LLMs, and can't explain the code it wrote.
Since even the engineers that know what's going on aren't actually reading all of the AI output any more (or, if they are, they're not keeping up with their peer's output), why would you care? I don't think humans should waste time trying to understand their code, it's too slow and costly, and the understanding will be blown away the next time the AI changes it anyways.
Software engineering is becoming pasting in vague-ish descriptions of what you want, and then manually testing that what the AI developed is close enough. It seems like math can go in the same direction too, with useful results that improve our technology getting put into a database for other AIs to consume. Removing humans from the loop can speed things up, especially as AI improves, especially when it reaches a self-improvement loop.
As I keep saying, software is no longer skilled labor.
This is a big if, right? AI can still generate subtle or even silly mistakes that any normal human, let alone a mathematician, wouldn't make. Besides, math is more than just getting a conclusion but to understand and to generalize new ways of solving problems. After all, mathematicians are a curious bunch. To quote Hilbert's epitaph: We must know. We shall know.
AI doesn't have to implement Hilbert's vision and be able to prove everything. I just has to out-prove human mathematicians.
Maybe we'll have some hobbyist dabblers, but any real progress will be done by machines that skip the human.
I don't think you can be coherently pro-AI without thinking that the human brain will be obsolete, unless you believe in some inherent magic that the brain is imbued with. The only other option is that you haven't thought through the long term consequences of the innovation.
Most research mathematics is pure mathematics which is completely useless. No routing algorithms. It's only relevant because we (or at least mathematicians) are interested in it. So an AI producing incomprehensible proofs would be completely pointless. That's why Tao insists on the importance of human understanding.
This is also the strategy I use for editing drafts of my books. I bring a printed draft to someplace nice (e.g. coffee shop or park) and read it all carefully, then I transfer the edits back to the .tex sources. I do several passes of this, until I feel the text + explanations are solid.
Reading on screen just isn't the same...
It's the old cliche of "if you only have a hammer every problem looks like a nail". Let's not fall into the trap of thinking that our life needs to be 100% about AI or completely devoid of AI. We can really use this thing to make our lives better.
Instead of wasting time on the question of whether we should use it, let's focus on HOW we'll use it.
And one thing about Tao: it's really refreshing to have an influential genius "around" who isn't a egomaniacal psychopath trying to rule the world through their XYZ corporation but, instead, being a reasonable and well-balanced person. Big fan.
The only thing to do is to be all in, or get run over.
Understanding was critical for the field to progress when only humans were involved but if humans are not needed to make progress, I wonder if we split into two worlds: an AI math-world where amazing new results continue at a rapid pace bottlenecked only by compute/cost and a human math-world where we understand a subset of the AI math-world as a hobby (similar to Stockfish vs human chess).
I am wary of AI in all aspects I am seeing it in but in many ways in mathematics seems to me the least troubling. It will change things in and the field will not be the same. Blacksmithing has not really gone away. You can still work as a farrier, if you like that sort of things. The tools that replaced a man working over a forge with a big hammer are part of a giant industry that is still producing works for the modern world.
The Busy Beaver game has lead to a better understanding of complexity theory and automata. Also, direct "hands on" work on improving proof assistants and related tools.
Btw, for those who are curious, the Busy Beaver Challenge wiki is a treasure trove of rabbit holes and curiosities:
And the term "artificial intelligence (AI)" has been the name of the field for 70 years and counting. If anything, "LLM" is a misnomer that's been lingering around since 2018-19. When the term was coined, these systems were relatively small, experimental, and could only produce impractical facsimiles of the English language. This is obviously no longer the case today.
It's been used to talk about computers playing chess, then machine learning, and now LLM-based systems.