> From the two postulates, Einstein derived the Lorentz trans- formation ...
If Einstein derived them, who is "Lorentz"?
The groundwork for Special Relativity was the study of electrodynamics and symmetries of Maxwell equations. The Einsteins paper was literally called "On the Electrodynamics of Moving Bodies" and never cites Michelson and Morley.
Einstein is not the first great scientist who are in denial of other important prior contributions, and he also not the last one. Newton also probably knew too well about Al-Haytham (Alhazen), arguably the father of modern science, and his breakthrough experiments but never directly cited Alhazen's works in his seminal books on Optics.
[1] Millikan, Einstein, and the Birth of Relativity (4 letters):
https://www.aps.org/archives/publications/apsnews/200403/let...
"Abraham Pais, who knew Einstein well and wrote his scientific biography, was certain that Einstein did know about Michelson's experiment before 1905. He points out that Einstein was over seventy and in poor health when he spoke to Shankland; at the first interview he probably did not remember that Michelson's experiment is discussed in Lorentz's 1895 monograph, the famous "Versuch", which Einstein had definitely read before 1905."
Science rightfully recognizes the mind that doesn't just first describe the idea but provides a robust framework to test it and communicate it.
> If Einstein derived them, who is "Lorentz"?
You can (re-) derive a lot of existing stuff.
Einstein was aware of Lorentz and the transform. He was aware of Poincaré as well. He knew the state of the art for his time.
Imho there's also a clean argument against the existence of LLM-understanding: chatbots have been unable to summarize to experts (see my reply to you in the other thread) their own findings.
Even after prolonged interrogation. They were unable to _compress_ their own findings. Thus they might not actually understand what they have actually done. (They might barely pass an oral thesis defense)
(One may object-- that proofs aren't data that can be "compressed". But then doesn't the process of abduction generalise the very idea of data? to.. ?)
But I'd like more explanation as to why Maxwell or Newton are any less replaceable: all 4 of "Maxwell's equations" have other names attached to them (the exception is Ampere-Maxwell law [1] which Maxwell contributed an important term to), and Newton had a number of contemporaries who were making similar discoveries but are often forgotten.
[1]: https://en.wikipedia.org/wiki/Amp%C3%A8re's_circuital_law
fair, there's lots of really awesome scientists, many of which not talked about even 1/100th as often as einstein
>when he doesn’t even belong in that conversation.
suddenly, the pendulum has swung way too far in the other direction.
What a depressingly peculiar time to be alive.
Reddit is over that way, my friend. You might find the crowd there more amenable to this nonsense.
Einstein’s fame isn’t the problem. It’s the narrative that he is somehow the greatest scientist ever, when it’s just transparently false when you look at his actual contributions and the historical momentum of the fields he contributed to. There’s nothing wrong with his fame, the problem is that it overshadows actual juggernauts, like Maxwell and Newton. Maxwell not being famous at all among the ordinary public is a great tragedy, when he genuinely is in the conversation as having have been the most important physicist to have ever lived.
> A few reflections on my "LLMs Can’t Jump" paper:
> My position paper recently got some traction here, so I wanted to share a few thoughts and clarify a few things.
> First things first: some people are framing this as "DeepMind is throwing cold water on AI for science" or claiming the paper argues LLMs can never make real scientific discoveries. This is NOT the case.
> This is a personal position paper, not the company's view on AI for science. This is also not my position. As a core contributor to AlphaProof (the first AI system to win an IMO medal), I know firsthand that my colleagues at DeepMind, other frontier labs, and academia have made amazing discoveries with LLMs and will continue to do so. This paper is NOT an "LLMs are a dead end" kind of thing.
> Rather, the paper is the result of a deep dive I took to study the invention of General Relativity. I wanted to explore what it would take for a modern AI system to make that exact kind of jump. Specifically, I focused on the equivalence principle—a key axiom that Einstein formulated through thought experiments grounded in his physical intuition. I was trying to figure out what it would take to give modern AI systems that sort of thinking.
> Giving AI this specific capability isn't necessarily the most urgent thing to do next. It is very likely that improving our current recipes will lead to many exciting discoveries in the near future. In fact, that is what I am personally working on these days (sorry to disappoint you!). It is also quite possible that I am wrong, and that simply scaling our current systems will lead to new inventions in physics and elsewhere.
> Nevertheless, this was my position last winter when I wrote the paper, and I'm sticking to it. I think that there are a few interesting ideas to explore in this space which could influence the next generation of AI systems. I was very lucky to receive a lot of interesting feedback about this position—thank you for all the messages!
It's weird because the equivalence principle is very unintuitive. Aristotle's Mechanics does not have it. It took almost two thousand years to discover inertia that is the most simple version of the equivalence principle. Einstein understood the idea of the the equivalence principle because he had a physics degree, not because he feel that in real life.
Moreover, if you ever have to study or teach Quantum Mechanics, physical intuition gets in the way. A lot of properties contradict the physical intuition but after a while you get use to them. If we continue with Einstein, the photoelectric effect does not aperar in real life.
Of course it does. How do you think your phone camera works?
(Note that the HP Chipmunk 9836 also had a scroll selector wheel in 84)
Most people are UI bigots. Once they get used to a first something, they expect everything to work that way and hate learning anew. They get stuck on keyboards, mice, trackpoint nubs, trackpads, trackballs, scroll wheels, or touchscreens and refuse to move on. Of course there are 'objective' performance tests for each including Fitt's test of accuracy and latency, as well as, cognitive load. I guess once you have a hammer, every screw looks like a nail.
So I'd be a bit careful with the "of course!". It may well be obvious only because that is the first mental model you latch onto.
1. Read the last sentence of the abstract, and
2. Reflect that frontier reasoning agents already increasingly integrate multimodal models.
I simply point out here, that fully accepting the paper’s premise, the paper’s conclusion isn’t limiting on frontier AI reasoning agents. The paper posits the necessity of multimodal world models and the limitations of LLMs. Frontier agents aren’t simply LLMs and do increasingly integrate increasingly capable multimodal models.
> Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes.
> Comments should get more thoughtful and substantive, not less, as a topic gets more divisive.
> When disagreeing, please reply to the argument instead of calling names. "That is idiotic; 1 + 1 is 2, not 3" can be shortened to "1 + 1 is 2, not 3."
TFA was actually about leaps of intuition, sadly.
One of the experiments I've heard proposed around here is to somehow create an LLM from all text up to 1980 or 1990 and see if it can get back to making itself.
Could be, but preventing leakage from more modern stuff can be challenging.
This was attempted with Victorian public domain content: https://www.estragon.news/mr-chatterbox-or-the-modern-promet...
I can't find the citation right now, but I think people found it was leaking anachronisms? So this probably wasn't as well filtered as the creator had hoped?
At a minimum, yes. IIRC, the sum total of all compute manufactured over history only reached the minimum needed to train an OK LMM in the mid 00s.
> How much could it infer from it?
Only way to find out is to try.
This is an interesting experiment but I wonder if it would be possible to prevent some sort of retrospective bias. For example, I’d expect the experiments that lead to relativity to be over-represented in our catalogue of scientific literature prior to 1900, just because in retrospect they were important, so the records about them were preserved.It would have to be a very intentionally constructed corpus, I think.
IMO, the mechanism isn't the important thing, the behaviour is. If you look at the step-by-step, we are also looking for the next word or motor action (and for whoever is about to suggest that we humans plan ahead, Transformer-based LLMs have been shown to also do this); as this is not a useful description of what it means to be a living brain, I'd say it's also not a useful description of what makes everything post-InstructGPT different from what came before.
I know the answer: because it leads to model collapse. But why is that? Wouldn't a smart model not collapse? It's seeming like they keep getting smarter because we keep pouring more of our own knowledge into them, not because they are actually getting smarter. And yes, sometimes a dumb but persistent bruteforcer can make new discoveries.
and i think this is exactly the crux;
the really big models need really big datasets
and current gen LLMs get a lot of training data beyond "all books + all of the internet"
the objection is then that producing this additional data would already confound it with pre "virtual cutoff date" knowledge (since the training data probably implies mathematical and SWE concepts that were developed post "virtual cutoff date")
But to prevent model collapse you need a way to pump down the entropy. Much like in thermo, it's an expensive and slow process.
My feeling is that a prompt would have to provide a vague description of a program that meaningfully passes something like a Turing test, an API to conform to, an expectation of novel construction (no 'ifs all the way down'), and then a requirement to search broadly and pursue promising ideas and not get hung up on the philosophy. Anything more precise feels like it would corrupt the test, but as it is that description feels doomed to loop before even trying the interesting parts.
Chessboxing was the invention of comics book artist Enki Bilal (and he's credited with this in Wikipedia). I first saw it in his Nikopol trilogy. Because life is weird, it then became a real thing.
It's unrelated to computers playing chess. It predates Kasparov's first defeat by Deep Blue. I don't remember any mention of computers being good at chess in the trilogy, either. Or any computers, for that matter.
It's actually possible to answer this question rigorously:
1. Define a scientific result which qualifies as a "jump". They should be frequent enough that they happen every year - otherwise one might say humans can't jump either.
2. Identify all such "jumps" in articles published in 2026, and use LLM with 2025 knowledge cut-off to re-derive these results with minimal amount of information.
It really irks me that people boost these low-effort articles just because they confirm pre-conceived notion that LLMs are limited
Until LLMs have some 0% error humans will have to be in the loop (even if they only serve to take responsibility of the process).
The only way to justify trillion dollar valuations is to sell cruelty-free robot slaves which replace white collar labor for pennies on the dollar.
Even though it's a complete and total sham, there's no American agency willing to or capable of prosecuting these firms for fraud, so they really have very strong incentives to keep up the lie.
Also possible if you make god
Saying you want to make workers more productive and provide better tools for people just isn't that sexy.
An LLM in isolation from its environment might as well be a brain in a vat in some dark cave. You need an external environment to sample from and act upon to make forward progress.
I see similar thinking in stories of how humanity got here. Religion has thousands of years adapting to this problem, every time we explain something, the goal post moves. Catholics today accept evolution (or least the church does), but it is the "jump" from monkeys to humans where God is the only explanation.
Just 5 years ago we didn't have a technology that knows more about everything than even most experts. We keep coming up with benchmark after benchmark and LLM/AI keeps destroying them. Now we've moved the benchmark to "the jump". Again, maybe it's LLMs or the way we currently do them that can't do this, but eventually something will.
We still don't have that. LLMs have shown time and time again that they don't know a single thing and are incapable of reasoning.
And we still don’t. What we have are simply very advanced search results aggregators with delusions of personality. Just because your fridge says “I” doesn’t mean it is a person.
"In abstract domains such as Mathematics or Computer Science, the Sense Experience (E) may be grounded in high-dimensional topology or have other goals such as generality or minimality."
But if such sense experience is possible in abstract domains via some high-dimensional topology, why could a sufficiently advanced LLM not develop an equivalent high-dimensional topology for domains like physics and use it to make creative leaps?
But we have no idea at how good humans are at that. Given the appalling failures of humans to handle even basic statistical situations like identifying that the same thing happens over and over, it might be that they are hilariously bad at creative leaps in abstract fields, it is just we have had nothing better available to measure against. We've spent about as long as decision theory existed trying to convince people to use it instead of flailing. Limited success, usually in exceptional cases.
And the paper seems a bit dodgy, we have models created with sensory data available. No reason a LLM can't be trained on more sensory data than a human can accumulate in one lifetime. There is a lot of visual data on YouTube.
The case study they chose is literally Albert Einstein coming up with General Relativity, something most scientists of his time were not able to do
One interesting (albeit sad) area which might be related are humans who are never raised with a first language. They seem to never developer abstract reasoning and even seem to lose the ability to develop it later in life. This might indicate there is some 'real world senses' -> 'direct language' -> 'indirect language' -> 'abstract abduction' hierarchy that develops, perhaps related to more real world abductions as a necessary side chain to developing abstract ones.
One of the obvious problems with this is just how difficult we find it to study intelligence purely in humans. We are measure a LLMs by a yardstick that is already known broken, but maybe this is still the right path.
For example "..ARC captures the logical leap, it misses the manipulative component—the physical sensation and embodied simulation..." makes lots of assumptions on how such a discovery must occur, e.g. through "physical sensation and embodied simulation". Results matter, not the path there.
For example, quantization of energy, at the core of QM, wasn't discovered through "physical sensation and embodied simulation" at all. Planck simply found that if energy is quantized, then one obtained the observed black-body radiation spectrum. There was no "physical sensation and embodied simulation".
This is also tied to halucinations: it is something that humans do (for writing fiction, and for "jumps") - but what LLMs currently lack is intellectual honesty. Coming up with bullshit is fine (and in this context valuable) - the important bit is putting those ideas through some form of rigor, or just immediately turn around and admit to talking shit.
So I'd arge that hallucinations are what prevent LLMs from doing this in a useful way.
If we could make progress in that area, maybe CoT could gradually decrease as it approaches its limit, or maybe the LLM could control the temperature of the next token itself (how this would be trained, I have no idea).
LLMs are inherently probabilistic, and there's currently no mechanism for producing an orthogonal directional change in the path traced through a latent space which is also contextually relevant (landing on a punch line).
In other words, LLMs are fundamentally incapable of making intuitive/orthogonal leaps in context.
It might be possible to add this capability with a new architectural component like transformers, but specifically for making "left turns"/intuitive leaps.
"Probabilistic next word prediction" and "humor" sit about as far apart as "modulating airflow with meat flaps" and "humor" do. One is an interface through which an action is performed and the other is a highly abstract capability.
Would you claim that a podcast comedian is fundamentally incapable of being funny because all he ever does is wiggle the air with his throat meat flaps? Probably not.
Absolutely nothing about "probabilistic next word prediction" forbids "making intuitive/orthogonal leaps in context". The interface is expressive enough.
And empirically? The "sense of humor" in LLMs is yet another "a function of model scale" capability. GPT-4.5 was reportedly funnier than both GPT-4o and o1. Fable 5 is reportedly funnier than Opus 4.x. It's one of those ever-elusive "big model smell" signs that are hard to measure with anything other than vibes.
Under the "humor as an opposed social intelligence test" family of hypothesis, what "being funny" reflects is the funny guy's ability to model and predict you and your reactions. For the comedian to be able to make the audience laugh, he must know his audience well, model it accurately enough to be able to spot the "breaking points" of humor, things they'd find unexpected and clever and thus "funny", and then weave those things into the jokes.
Then, a bigger LLM gets better at humor because it has a more accurate model of how humans think of things - including the "ha-ha" gaps. It's a "theory of mind" capability. It's not "special", it's just hard.
I think you're saying that you can eventually train models to arrive at that destination by training on existing jokes, effectively encoding these leaps as probabilities.
In that case, the model isn't actually making an intuitive/comedic leap; they're just following new probability chains in attempting to approximate examples they've seen in training.
I'm suggesting that something architecturally different is necessary to create a model which can make intuitive/comedic leaps.
Try to get a frontier model to write a clever, funny joke which hasn't been seen before. Or, try to get it to make an intuitive leap that leads to a novel discovery.
You can use them to guide your own efforts along these lines, as a sounding board. But with current architecture I just don't think either is possible for an LLM to do on its own.
This is what I refer to when it comes to larger models like Fable 5 being funnier. They are more capable of doing that. They can deliver that "sudden orthogonal leap from context" of yours more reliably.
It's not a "fundamental inability" and never was. If you crank the scale up and a capability appears, "current architecture" was never the problem.
https://news.ycombinator.com/item?id=49136070 https://news.ycombinator.com/item?id=49096837 https://news.ycombinator.com/item?id=46890333 https://news.ycombinator.com/item?id=46870562
All of these are titled “LLMs Can’t Jump”
For this paper specifically, after reading the abstract [2], I felt almost certain that the author would have used Judea Pearl's ladder of causation (https://web.cs.ucla.edu/~kaoru/3-layer-causal-hierarchy.pdf) but they did not. Would have probably been a better argument to make.
[1] paper in quotes because it may never get published (it is over 20 pages atm). the core argument is that lack of native adjacency resolution makes problems harder and sample inefficient, not impossible
[2] "Using Einstein’s formulation of General Relativity as a case study, we demonstrate that LLMs are structurally incapable of creating new foundational axioms, particularly when observational data is scarce. "
Also, the claim that 'LLMs are structurally incapable of creating new foundational axioms' is provably false depending on where you place 'fundamental'.
In math its simple: does the verification say its okay.
If its mechanical: is any property better than what we have already.
etc.
Every "can't" of this nature was followed by a discovery of "they can, just poorly", and then by that "poorly" improving steadily generation to generation.
The paper doesn't provide a way to measure or quantify this elusive "jumping" capability, not even as an approximation. It just throws "can't jump" out there, as if "abduction" is an established class of problem with known computational properties and requirements that the LLM architecture fails to satisfy. It's none of those things - and the paper makes the claim without backing it by anything but rhetoric attempts at persuasion.
The proposed solution is also dubious. The empirical track record of dedicated "world models" for reasoning and problem-solving is, frankly, downright abysmal. Even integrating multimodal data into LLMs has failed to yield general reasoning capability gains.
LeCun's misadventures in the field aside, the main frontier lab that pushes in favor of "improving reasoning via multimodal fusion" is GDM - and Gemini isn't exactly a paragon of frontier reasoning capabilities. It has strong multimodal capabilities, but lags behind both OpenAI and Anthropic in performance outside that - while Anthropic is the lab that always treated multimodal grounding as an afterthought, and still trades blows with OpenAI at the very edge of the performance frontier. Multimodal grounding seems to work great as a way to improve an AI's ability to deal with those specific modalities, but it falters outside that.
Now, it's not impossible that everyone who tried multimodal world models for reasoning is just doing it wrong, and there is an undiscovered recipe for multimodal grounding that results in a step change in AI capabilities. But the results we have so far suggest it to be unlikely.
My opinion of claims like "LLMs need memory to manage codebases" has also hit the dumpster bin a while ago.
Why would knowing how to make a maintainable change to a codebase require any more "memory" than knowing how to play an optimal chess move? The codebase is the memory. A sufficiently capable LLM can ingest it, figure out what changes to make, and make them.
Were that the case LLMs would’ve been phenomenal code monkeys from the get go. They were not. They still are not.
What LLMs are "fundamentally incapable" of doing has striking parallels to https://en.wikipedia.org/wiki/God_of_the_gaps
I just don't see any of the LLM users around at all. Clearly some force is guiding them all away from thinking any of the "leap of faith" thoughts that I am thinking.
You may think this is not a good test because an older (or say a smaller) LLM can study from the knowledge on the Internet and build. But we are like that - we can access the Universe through our senses.
Can we ever produce anything that is beyond this Universe? I think an LLM that is lacking in knowledge can build more complex systems as long as it can access more data.
Turns out temperature is pretty bad too, you can find ways to sample from deeper in the distribution without distorting it. Great example is XTC (exclude top choices), In a few weeks/months it'll also have a proper scholarly paper with peer review.
And then such jump can be verified by a machine so human kind of plugs the intelligence gap.
That’s pretty exciting.
I always liked to provocatively call LLM „the new calculator”. Calculator for language.
We are so focused on creating a standalone intelligence that we didn’t notice how we massively augmented our own. That could be considered transhumanism holy grail if only interface brain-LLM was faster.
People need to understand that these things are tools. And every tool needs an operator to function. Tool doesn’t have its own goals, needs, wants or motives. It won’t do anything out of its own, it always exists in context of someone telling it what to do.
In light of that most of the panic and fear mongering is rather ridiculous. Calculator won’t replace you. It wont take over the world. It is just a tool.
You write a book with book generator? Cool, it can be used for this. We will judge output, not the methods. Sometimes we will judge people who have no taste in literature.
But then , one difference I found on LLM's is on scaling laws, where at some point, it have interesting emerging properties, that nobody thought would be possible now being possible.
Tools: Though It is a tool, but powerfull tool that is automating existing manually done jobs, large scale. Adapting to new roles where we definite goals, needs and motives judging of AI output, at large level is an issue. And humans are used to day to day repeating job, doing same thing repeatedly. Now the AI is taking over that. I see that is a challenge.
Also I see we are moving to creative world, where we will spend more time on creating something really new, leaving mechanized parts to AI.
The only way a LLM can come up with new ideas if the "idea" appeared as a generalisation durring training or if it was achieved using reason in chain of thought.
A jump in intuition comes from automatic processes reorganising the relational structure of conceptual models. There is no reorganisation of the model durring inference.
one could argue that model can reorganize / interact with prior knowledge captured in text form (edit files) hence there can be reorganisation