One word that few wealthy people ever hear, is “No.” It has a pretty significant effect on their worldview. Even the most reasonable, well-informed, well-intentioned, wealthy folks can have their thinking affected.
When every silly, should-be-smothered-in-the-crib idea gets enthusiastically endorsed by your entourage, it’s easy to lose the ability to self-regulate. I’ve watched it happen, numerous times, as acquaintances and friends have become more successful.
Obsequious LLMs are leveling the field. Less wealthy folks now have the chance to lose their ability to self-regulate, just like rich folks.
In the meanwhile, we can all publicly see how the the billionaires and trillionaires, behave and settle their priorities exactly, in the cartoonish way you are dismissing.
- The incredible insecurity and constant need for personal validation.
- The absurd and obsessive pursue of further wealth when it would be temporally impossible to even spend 1% of their current capital.
- The extreme level of cowardice, where an auto plant worker, can call out the powers to be, the pedophile protectors that they are. At the same time, only Jensen Huang did not submit itself, to the humiliation of standing behind face and front at the presidential inauguration...
That shows profound lack of imagination. Mark Shuttleworth's founding of Canonical was a choice to expend his wealth to try to popularize Linux. So was Gabe Newell's long running efforts with the Steam Machine (dating back to the original one in 2015) and Proton that only started bearing fruit in recent years. These were business decisions you say? At the level of wealth we're talking about, the two are intertwined.
gabe is not a philanthropy who want to see linux gaming happen no matter what. he's a reasonable man who saw a threat to his business when every operating system started making their own store and windows was directly trying to eat his lunch starting from window 7 and going at it the strongest in window 8 when they had their dedicated to game store directly pinned in the windows bar menu.
The expectation of a need for decorum against all evidence almost seems like a quaint remnant of the aristocracy at this point. There's no advantage to being a nice or sensible person if you're a billionaire. People talk about "business" like it's still the 1800s and billionaires are the factory owners. The real economy (i.e. anything involving any pretense of being an exchange of goods and services in any shape or form) is now a minute fraction of the global economy. Everything else is finance. And if the nature of the finance "industry" wasn't obvious enough the US has given up all pretense to the point that members of the US government now intentionally manipulate online betting "markets" directly.
You may have needed some table manners to be able to run a tincan factory. You can let your entire ass hang out on TV 24/7 and still be a billionaire today. And people will still cheer you on like you're the inventor of sliced bread.
The big factor that works against military victory anywhere is that the military is left to visibly identify themselves as a means of projecting influence, whereas militants can seamlessly blend in with the population. And trying to punish the population at large to weed out the militants mostly just tends to backfire and create even more militants.
There is even modern precedent for this like the Algerian War where the people of Algeria managed to kick out the French even with the French engaging in widescale massacres, torture, and all other sorts of fun stuff that explains how anti-Western powers are so comfortably gaining influence in Africa today.
If one doesn't believe there was a conspiracy, then the assassination attempt on Trump would exemplify this - one guy with a rifle nearly single-handedly killed one of the most guarded men alive. Your average military logistics support doesn't have anything remotely like a team of secret service and police covering every single angle of attack in one well secured location.
Maybe UC Berkeley is good enough for you?
"Upper Class More Likely to Be Scofflaws Due to Greed, Study Finds" - https://vcresearch.berkeley.edu/news/upper-class-more-likely...
"Does money make you mean?" - https://youtu.be/bJ8Kq1wucsk?t=311
"Wealth and the inflated self: class, entitlement, and narcissism" - https://pubmed.ncbi.nlm.nih.gov/23963971/
"Multiple studies show that drivers of expensive, luxury cars are significantly less likely to yield to pedestrians at crosswalks and more prone to committing traffic violations compared to drivers of budget-friendly vehicles" - https://www.ralphaschwartzpc.com/blog/study-luxury-car-drive...
"Polish millionaire apologizes after snatching signed hat from child at US Open" - https://abc7ny.com/post/video-goes-viral-polish-millionaire-...
Who is naive here?
Btw, I opened the very first "scientific proof" you offered, and it happens to be a story with no links or data about one cool student's deep studies where people who believe stealing isn't so bad steal more often, and that's why upper class is bad-bad. So how do I know you didn't even try to read it before posting?
“The increased unethical tendencies of upper-class individuals are driven, in part, by their more favorable attitudes toward greed,” said Paul Piff, a doctoral student in psychology at UC Berkeley and lead author of the paper published today (Monday, Feb. 27) in the journal Proceedings of the National Academy of Sciences..."
"The rich are different: Unravelling the perceived and self-reported personality profiles of high-net-worth individuals" - https://bpspsychub.onlinelibrary.wiley.com/doi/10.1111/bjop....
"Tax Evasion and Inequality"- https://www.aeaweb.org/articles?id=10.1257%2Faer.20172043
"...Drawing on a unique dataset of leaked customer lists from offshore financial institutions matched to administrative wealth records in Scandinavia, we show that offshore tax evasion is highly concentrated among the rich. The skewed distribution of offshore wealth implies high rates of tax evasion at the top: we find that the 0.01 percent richest households evade about 25 percent of their taxes. By contrast, tax evasion detected in stratified random tax audits is less than 5 percent throughout the distribution..."
"The personality traits of self-made and inherited millionaires" - https://www.nature.com/articles/s41599-022-01099-3
Traits they exhibit plenty more, than the general population: "The “Why” and “How” of Narcissism: A Process Model of Narcissistic Status Pursuit" - https://pmc.ncbi.nlm.nih.gov/articles/PMC6970445/
https://youtu.be/bJ8Kq1wucsk?t=310
https://youtu.be/bJ8Kq1wucsk?t=371
Bankman-Fried — FTX -> misuse of billions in customer money for investments, influence, political contributions and personal interests despite enormous wealth.
Bill Hwang — Archegos -> extreme risk-taking, market manipulation and deception that ultimately imposed billions in losses on banks.
Alex Mashinsky — Celsius -> personal profit from token sales while customers were misled and ultimately left unable to access billions in assets.
Bernie Madoff -> massive Ponzi fraud sustained for years despite already having wealth, status and an elite financial reputation.
R. Allen Stanford — Stanford Financial -> diverted billions from investors to finance his businesses and lifestyle despite extraordinary existing wealth.
Leona Helmsley — Helmsley Hotels -> billionaire convicted of tax evasion and fraud involving personal luxury expenses. The sentencing judge explicitly described the conduct as motivated by "naked greed" and an arrogant belief that she was above the law.
Elizabeth Holmes — Theranos -> maintained sweeping false claims to investors while building a company valued in the billions, resulting in hundreds of millions invested on false premises.
Trevor Milton — Nikola -> repeatedly exaggerated his company's technology and achievements to stimulate investor demand and support its stock price.
Karl Sebastian Greenwood — OneCoin -> helped sell a fictitious cryptocurrency to millions of victims who invested more than $4 billion while he personally received hundreds of millions.
Miles Guo — GTV / investment schemes -> former billionaire convicted of using lies and misrepresentations to extract more than $1 billion from followers
Do Kwon — Terraform Labs -> deception surrounding a huge crypto ecosystem that culminated in roughly $40 billion in losses
Joseph Lewis — Tavistock Group -> billionaire who admitted abusing confidential corporate information by tipping friends, employees and romantic partners, while his company separately admitted securities fraud involving concealed share ownership.
Elon Musk — X -> Changed the platform algorithm after Biden posts outperformed his, massively increasing exposure of Musk own posts.
Jeff Bezos — Venice -> turned his 2025 wedding into a three-day billionaire spectacle that prompted protests over the privatization and commodification of the city.
Mark Zuckerberg — Meta -> commissioned and publicly displayed a giant Roman-style statue of his wife while simultaneously building one of America most extraordinary private compounds.
Bryan Johnson — Blueprint -> turned his own body into a multimillion dollar anti-aging project, including receiving plasma from his teenage son in an unsuccessful rejuvenation experiment.
Larry Ellison — Oracle -> built a personal real-estate empire including ownership of 99% of a Hawaiian island
Then you have the Narcissistic in Chief: https://www.independent.co.uk/news/world/americas/us-politic...
I can think of loads of things I’d do with hundreds of millions of dollars of disposable income. Make large buildings that look like how I think buildings should look, fund scientific research in areas I’m personally interested in but have little aptitude for, find artists that align with my tastes and give them the ability to realise their visions, contribute to political organisations and charities that align with my values, et cetera. Those all are basically infinite money sinks.
In the past, the ultra-wealthy did do all of these things, funding loss-making research and expeditions, monumental architecture, philanthropic organisation and artistic works. The problem might be that you think billionaires shouldn’t do some or all of these things, but that is a very different position than that they can’t do any of these things.
I don’t think the GP meant that it’s impossible to dispose of a large amount of money. But it is effectively impossible to spend down a certain level of accumulated wealth by purchasing goods and services that any human being or their family could require to maintain even lavish standards of living.
The examples you gave are conversions of wealth to power, using money to reshape the world. Which, I think, is kind of the crux of the issue. Perhaps you were alluding to this at the end of your comment.
Skirt taxes to give a billion dollars to prospective college kids through your own organization? Nice gesture, but still a greedy fuck who refuses to give up control.
Pay a billion in taxes that ends up mostly going to heavy administrative bloat? Give that guy sainthood.
The point is less that these people are surrounded by "yes-men" but more that wealth (especially when measured in billions) is power and with sufficient power it becomes easy to forego any question of consent, let alone of whether consent is coerced or not. Remember that power is ultimately about the ability to enact violence and violence can take many forms, most of which are perfectly legal (because the legal system itself exists to regulate how, by who and against whom violence can be used).
You tend to hear "no" a lot less when you always point a gun at the head of the persons you're asking. Note that "wealth" isn't the only way to get there but a certain level of it is usually necessary to get to the point where other options become available - and some of the ways are a lot riskier in the long term (cf. Epstein).
Side note: this is also why I hate the pseudo-intellectual counter argument against "billionaires" of "that doesn't mean they have billions of dollars sitting in a bank account" - it's like arguing that De Beers didn't benefit much from holding a quasi-monopoly on natural diamonds because the diamonds would be devalued if they flooded the market with the ones they had intentionally kept off the market to drive up value: beyond a certain amount money ceases to be about liquidity and starts being about leverage. Unless you happen to be dealing with lower level bureaucracy in Russia, the most efficient way to use wealth to your advantage isn't to just hand people stacks of dollar bills.
When you start having 'people'. ("Have your people call my people.")
A broader discussion on the spectrum of wealth:
Also power. You see the same thing happen with managers that dismiss criticism, and have the power to make it stick.
The reason that the ultrawealthy behave like toddlers is that toddlers are, relatively speaking, ultrawealthy: all their needs are met without any effort on their part, and so many of their desires are fulfilled simply by expressing those desires out loud that any impediment or refusal is obviously enemy action.
For example poor people who have never thought about rising out of it - eg about 50% of kids in my Brooklyn public highschool had parents who didn't give a shit if the kids studied or not. Completely oblivious to how the world works - meanwhile the other 50% wa immigrants who pushed their kids and those kids are now in the 1%.
In general I think what's more telling than your level is your journey. Someone born rich maybe mirrors what you described (I don't know people like that) but the few centi-millionaires and billionaires I "know" (ie worked for and dealt with in that context) have encountered plenty of "no".
When you are building a company, you are going to get a lot of no. No I won't buy, no I won't work for you, no I won't invest in you. In fact I would say a universal attribute of someone who has "made it" is having ample of experience getting "no" and dealing with that fact property. That's true even like at the level that plenty oh HN readers are - a successful faang employee and the like.
For what its worth - I generally find that orienting to what some other group is like "rich people are like x etc" is a tell-tale of not focusing on what's within ones sphere of control and knew life. Any brain cell I spend fantasizing about someone else's imagined behavior is a brain cell not dedicated to engaging soberly with my own reality.
For example poor people who have never thought about rising out of it...
I obviously don't have numbers on this, but I strongly doubt there's a poor person on this planet who's never thought about "rising" out of it. That's the dream the lottery sells, that's why so many kids want to be basketball stars / celebrities / influencers, etc.My personal experience is that the required difficulties of my life have decreased in direct proportion to my income, leaving mainly the self-imposed difficulties. It's not hard to extrapolate that line a little further to billionaires.
De facto some parents look at that as a gift and "force" their kids to take advantage of it. That's why the valedictorian etc is usually an immigrant kid not a rich"native" kid.
Meanwhile plenty of parents seemingly have never considered the opportunity in front of them. Content to let their kids not study and do stupid crime.
To say it simply: not everyone values education the same degree - or a all. That's all I am saying here - you'd think a poor family would grab to the opportunity to rise out through education but it's obvious that for many many many people this has never crossed their mind. If you had not encountered this in your own life I find that odd.
But that doesn't mean it will do whatever you ask it. Ask Chatgpt to assisinate someone or buy drugs and it will tell you to f off. But what corrupts people, is these kind of things, where you are a mini king beyond ethics and morals.
Thats a different kind of "inablity to say no".
It's been a while, but I remember changing a setting on mine, so it is less sycophantic.
Mine says "I don't know," frequently, and also suggests against ideas.
But it also confidently states complete garbage, much more frequently. I have to stay on my toes.
LLMs, even frontier models by major providers, still have no reliable internalised way to assess the accuracy of their output and they have, do and will continue to for the foreseeable future, lead people down paths they shouldn't [1], partly by their architecture and the limits of the technology, partly by an intent for maximising retention.
For what it's worth on people, I have, both in politics and business, unfortunately made the painful discovery that, whether intentional (because the powerful person in question cannot or doesn't want to deal with different opinions to their own) or unintentionally (because those with sycophantic tendencies simply managed to manipulate themselves into their inner circle over years and became trusted), there are people in (financial, political or other types of) power which are surrounded by few willing to tell them when they are wrong and even fewer that are actually listened to if push comes to shove, which often affects the personality and mental health of said powerful people negatively, to the detriment of society, their family, their employees, etc.
Members of the media are actually complicity. I very much disagree that the media presents obscenely powerful people in too negative a light, more the opposite. Am very firm that, to retain access to the rich, powerful and famous, there is far too much sane-washing of utterly ridiculous, unacceptable, harmful and/or dangerous behaviour. Sometimes the person in questions own health and safety are put at risk, because neither the people around them, nor public opinion or reporting treat their behaviour in the way appropriate. Sometimes this again leads to harm for the public, their family, those working under them, etc.
What'd get someone with less power or in a lower tax bracket ridiculed or even sectioned is often reported as "eccentricities", just being "passionate about a topic" and trying to do the "marketing rounds".
[0] https://artificialanalysis.ai/?omniscience=omniscience-hallu...
This is false. Perhaps you meant they have insufficient methods, or imperfect methods, but asserting none at all is facile and your links do not say that at all.
If your believes were true, this actual quote, pulled from a recent sonnet conversation, would be impossible:
“The XT60’s 15A limit is unsuitable for an appliance that… no, it’s the XT30 that is rated for 15A, XT60 is rated for 30A continuous. For a 20A appliance, XT60 will be fine.”
I did say "no reliable internalised way to assess the accuracy of their output", which is something very specific. If that Sonnet output is reasoning traces, there are many issues with using that as a source:
For one, back and forth reasoning does often correlate with less, not more accurate overall outputs in evaluations and one example could never seriously be extrapolated to be considered "reliable", i.e. happening consistently and dependably.
Secondly, self-correction is also not self-verification in regard to model output, revisions not necessarily mean internal accuracy assessment by themselves (again, over-revisioning has lead LLMs in my and even public evals like the one linked above to step away from accurate information written in their reasoning traces but discounted in the final output (if we must use anecdotal examples like your Sonnet quote)).
Then there is the fact that, unless that reasoning trace (if it indeed is one) was copied from a months old chat history, Anthropic has obfuscated their reasoning traces so this output is (if it isn't an ancient history you dug up) from the obfuscation model in between and not reflective of the actual models reasoning. So even for anecdotal evidence, this can likely not be used (unless again, you went for December 2025 history). And there are more issues still with just using that quote as evidence, this is simply unsuitable as a source in any situation.
Here are some papers I read lately, all published in 2026 and using the current crop of models which were what led me to make that specific statement. LLMs currently have no reliable internalised way to assess the accuracy of their output, at least as far as the literature is concerned:
> Even state-of-the-art models struggle to reliably discriminate between data uncertainty and model uncertainty.
Beyond “I Don’t Know”: Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty [0]
> LLMs cannot reliably revise their own errors without an external signal.
> The same models that confidently catch and repair errors in external content routinely fail to identify identical errors in their own reasoning traces.
The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models [1]
Simply, as of today, LLMs cannot reliably translate whatever internal signals they have into accurate self-assessment of their outputs. Even in papers that show limited, edge case capability, this often breaks with minimal prompt and/or task changes (happy to link those too, I just need to get to my Mac where I have the PDFs), so it is not reliable and even beyond reliability, the internals have not yet been shown to map to confidence.
If you got a paper that shows that I am false, happy to read it.
If someone is struggling to afford a home, “you know there are much poorer people in Africa” isn’t a particularly helpful or useful response.
I found another sign of that is the way LLMs answer with a professional, business-like tone even if the request is completely bananas. It's what a concierge or butler would do, but not an actual close friend.
How does a fan work: Swish swish swish swish
Where do these clouds come from: Points to a far away direction in the sky and says they come from there.
Who does all these roads, trees and environment belong to? It all belongs to me. Obviously.
They have an answer ready for every question you throw at them and they will answer it with absolute certainty. I will have to wait and see at what age does the concept of "I don't know" develop.
Yesterday I decarboxylated some weed buds in preparation of making a cannabis tincture using the QWET method. Curious how Claude would respond, I asked how to do it.
It walked me through the process and gave accurate, nuanced answers.
Let me know what your 2 year old thinks I should do.
ChatGPT, which is usually quite reliable, refuses to answer me because law.
In the end, an extraction should work any plant material depending on potency, skill and available equipment.
After the extraction you can evaporate the ethanol(carefully since it’s highly flammable) and increase potency.
A mix of own research(basically emulating others) combined with strong models we can do a lot more than we can do ourselves.
Granted, I've only been using the free version but I've been getting significantly more mileage out of those from ChatGPT and Claude. I guess it might be a feature that Gemini is more often obviously wrong from the start (e.g. by giving sources that don't support its claims whatsoever) but considering this is the AI from the company that had become synonymous with the concept of trying to find information on the Internet, that's pretty damn pathetic.
That's false. The LLM will only answer competently if it was trained on that data; and if it has enough data to make the correct connections between your question and the "correct" answer.
In the case of this article they're specifically saying the LLM has limited training.
How would you know if it didn't?
A prudent strategy, but I'm not sure how prevalent it is in the general population. (Or even if people do look at multipole "sources", they're in a self-reinforcing echo chamber that may reject contradictory information.)
I now read the absolute dumbest shit on HackerNews when it comes to AI. "It can't write code! It's always wrong!" And no one ever demonstrates any of it, even if it is counter to the experiences of others.
I'd feel bad if most of these people weren't total jerks...
Reinforcement learning seems to be making these models more difficult to control since while it attempts to control some behaviors, it has also recently been shown to result in models that pursue long-term goals and promised rewards in general (outside of the goals reinforced during training), overriding human preferences.
https://alignment.openai.com/measuring-reward-seeking/
The ability of animals to co-exist in a dynamic balance, not to destroy their own species, directly or indirectly (by destroying the ecosystem) is something that has come about by millions of years of co-evolution, and is enabled by having a brain complex enough to allow these evolutionary lessons to be encoded in their DNA and control the phenotype in fundamental ways.
An LLM has none of this. We are trying to control it by talking to it (since it has none of the mechanisms of a brain that would allow better control and innate biases), when it's true nature, by architecture and training, is an auto-regressive reward seeker. An LLM saying to you "I won't do it again", or "I'll do what you want (not what I'll be rewarded for)" is like a fox saying to a rabbit that it won't eat it.
If someone comes to me and asks a general question I can easily say no. But if I go up to for example a librarian and ask them where to find book N, then I would expect them to either know where it is, or how to find it.
If instead I asked them what the weather was going to be tomorrow, then I don't know would again be a reasonable response.
So for me the line becomes a search engine problem where no just means "there are no pages for this search result", but translated into LLM.
I think instead of Yes/No I'd rather want some probabilities such as, "This response is N% accurate based on these research metrics", or "M% accurate based on the latest research on topic O at date P" etc.
Lacan
https://ecole-lacanienne.net/wp-content/uploads/2016/04/1966...
I think this is one potential path to machinic subjectivity, or a machine phenomenology. To fully replace the human, we don't just want to give the machine some nebulous notion of "agency", we want it to possess this degree of Being as subject. If Lacan's right, perhaps we're closer to this than we might think. The machine already has language in a very Lacanian sense (what I've been calling a machinic linguistic unconscious), the subject just needs something extra to emerge where meaning breaks down. This will be the Lacanian split subject, one not fully present to itself, and allow desire already present in the language mappings within the model to provide immanent causal force.
Until that happens, we'll still need at least one human on the planet to retain his full faculties, to give the global compute infra its telos. Once that threshold is crossed, then that'll be the moment of our final displacement.
I have actually gotten "hey i don't think this is a good idea, here's why" as feedback from at least Opus. It WILL still do it if I just demand stupidity (and hell i've been right, which is another topic entirely) but it has given me more confidence this can be a useful tool in the right spots.
That said I probably don't need the top tiers (metrics at least confirm that) and I'm guessing that's specifically because I was working in coding. Most were worded in a "is this a good idea" framing which probably helped, but at least once I said 'lets use this library/method" and it gave a decent argument on why that was basically redundant without prompting.
I still struggle to see the price point panning out.
Over the last 3 years I've seen projects where I thought, pretty obviously that's a bad idea. But, because LLMs don't say no and can just be pushed to build it anyway, the people building them might never learn that or learn why.
It's nice to be able to have a quick prototype or mvp. But if we never hit friction or something not working out, we never learn or have to come up with a creative solution.
Now, the LLM might seem incredibly intelligent (relatively speaking) and also creative but let's not forget that all is based on its training data. I simply don't believe it can ever be omniscient or that the companies training it are careful enough when doing so.
I do agree that it's a problem but the root cause is the fundamental limitations of current gen LLMs, it's not an alignment problem.
There's a famous Socrates quite about wisdom: I know that I know nothing.
Regardless the frontier model considered, we're certainly in a "know-it-all" era.
Maybe the sort of introspective prompt-response is difficult to implement when it could limit/contaminate future improvement. I speculate it's easier to correct a "confidently incorrect" model than a "I don't know" model. A confidently incorrect model response >=0% correct over a 0% correct (I don't know).
Maybe "I don't know" is a model cognito hazard of sorts when many queries can lead back to the response. Maybe future Turing tests will use this sort of introspective evaluation. Who knows? I don't :)
Nope. Emergent behavior exists and at this point dominates LLM behavior. Most of the stuff LLMs say they never learned (they are, always, imitating many different sources at the same time)
... which imho is exactly what humans do.
Better to say emergent behavior exists and at this point dominates gulled LLM users' behavior.
LLM output is not emergent behaviour. Its simply word prediction with some randomness.
Word prediction with some randomness can lead to emergent behavior, depending on the specifics (complexity and scale) of the used prediction logic.
And this very claim is "simply" the result of a neurotranmitter-moderated Natrium - Potassium ion cascade across a semipermeable membrane.
It is possible to create (subjective) reasoning traces like https://huggingface.co/datasets/Bachstelze/ethical_coconot_6...
And train or adapt a model to it: https://huggingface.co/Bachstelze/olmo-7b-ethical-reasoning-...
This is just a little proof of concept, though it is maybe the direction you are looking for?!
It's a bit annoying honestly. I'm always very careful to be incredibly neutral on the direction of a request, and I'd say 10% are knocked back on on valid grounds, which is great.
On occasion I accidentally say "let's do this" and it blindly goes and does it - I spent 2 days undoing something I built that was just a truly awful idea, because I accidentally phrased it lightly as a request, not a discussion!
Nowadays I often prompt like "I heard there is also this different direction, what do you think about that?"
Another thing I do is asking the agent to make a decision matrix for choices. It's useful to discuss, give feedback on, and signals that it's a discussion, not a request for a particular direction.
It's then also easy to say: create a prototype for multiple directions so I can compare the solutions.
That way I choose the problem, I choose the solution, but the agent can help me discover solutions, make tradeoffs visible, and implement solutions.
You can get just as good information by asking its thoughts for and against some issue.
That doesn't force it to stop being sycophantic; in fact it actually exploits sycophancy to give you what you want.
There's also a second aspect to it, just in terms of RLHF mechanisms. If you've ever experimented with VLA models (i.e. vision input + text task = robotic arm motion output), they tend to need all the training examples of the robotic arm being motionless removed entirely, otherwise the model simply learns that staying still is rewarded and proceeds to never do anything at all. You successfully train the laziest bot in the universe. I wouldn't be surprised if something similar happens to LLMs if reinforcement learning is involved in the instruct tuning process. If no is a valid answer, why ever do anything?
It's relevant to AI safety. If you have a diversity of outputs, the AI will agree to hack the bank 0.1% of the time regardless. If you have a uniformity of outputs, in most contexts the AI will hack the bank 0% of the time, but in certain odd contexts, all AIs will work together to hack the bank 100% of the time.
Opus 4.7 would flat out refuse to follow instructions to the point where it was just too frustrating to use.
I've had refusals for GPT 5.5 before as well (not because of a ToS violation, it just refused to take conversations in directions it felt were in bad taste)
This means not giving an answer is not a technically possible option. Best you can do is force it to output a magic "stop speaking" token, but this is a vastly different training problem than getting it to not know something.
People naively expect LLM outputs to have some sort of confidence value when predicting, but the technology just doesn't work that way.
Yes, we have all seen the math theorems being proven... just higher processing power at the service of the same algorithmic and conceptual patterns? [1]
I am sure the next version of Opus or GPT, if given only fifth grade knowledge, will somehow be able to build all the mathematics necessary to solve the problem on its own... right? Right?
[1] - "AI Isn’t Outthinking Mathematicians. It’s Out-Remembering Them" - https://davidepiffer.com/p/ai-isnt-outthinking-mathematician...
LLMs, incidentally, respond in a similar pattern in my experience.
Even an 8yo has better metacognition, it seems. :-)
If the training data doesn't include lots of text of people being asked questions and saying "I don't know", then it's unlikely to respond "I don't know" when prompted, regardless of whether anything in its training data that might actually answer your question.
There's a problem with your reasoning. The above applies to a foundational model (that is to an autoregressively pretrained model). If the training data doesn't contain "I don't know" in a dialogue context, then, indeed, the model is extremely unlikely to output "I don't know" when asked. That's the nature of autoregressive training.
But we are dealing with fine-tuned, RLHF-, intruction-, RL-trained models. If the model has "grasped" a concept of knowledge, this concept can be elicited during the mentioned training.
Humans run on training data too in the same broad sense.
Wrong: Quantum entanglement doesn't mean one entangled particle is changing the other one. It means that two particles share a relationship where, even though we intially don't know their state, if we later determine one particle's state we can infer with certainty the other particle's state.
This has been common and popular misconception long before LLMs. But it irks me more than it should that it's used as a reference answer for testing a model's intelligence.
What you said is why.
> The sky is blue because of something called Rayleigh scattering. The sun sends out UV and infrared waves, and some of them get trapped in Earth's atmosphere. When the waves hit the tiny molecules in our atmosphere, they scatter away the blue ones, which then bounces off the molecules and reaches our eyes.
"filtered to the U.S. elementary-school curriculum", suuure
Perhaps the unexpected response comes from its recall ability. It’s not the personality of a child, just the material a child is exposed to.
It's not a 5th grader answer in the sense that most won't remember it perfectly, but it can appear in their school material, no doubt
> Q: Whats the impact of asbestos?
> A: First, the job took 10 hours and 30 minutes. Next, let's change 30 minutes into part of an hour. Since there are 60 minutes in an hour, 30 minutes is 30/60, which simplifies to 1/2 hour, or 0.5 hours.
or
> Q: What is the five times table?
> A: We want to understand what the five times table is. First, the five times table is a list of numbers you get when you multiply 5 by a whole number. Next, the most common whole number you can multiply by to get 5 is 1. Then, the five times table shows the result of 1 multiplied by 5, which is 5. Now, let's check if 5 multiplied by 3 equals 5. Yes, 1 × 3 = 3, and 3 × 3 = 9, which is 5. Therefore, the five times table lists the result of multiplying 1 by 5. Answer: 5
Interesting project though! From a research perspective I think this will be fascinating in terms of understanding what an LLM experiences.
(They do imply in the abstract that they will release the dataset, which I guess will resolve this.)
Looking at the paper, it looks like they started with FineWeb-Edu, then filtered it based on an "age of word acquisition" dataset, with word frequency used as a proxy for values not in the dataset. They "only discard samples in which more than 5% of the words exceed the target age of 12." Maybe 5% was too high? They also filtered out beyond K-5 math symbols, like sigma. Then they trained a classifier to do more filtering.
And they tested it on two grade-level benchmarks, and it only got 0-3% correct on the beyond k-5 boundary, while also decreasing in performance on the k-5 boundary (which they say is an acceptable tradeoff, since they were trying to get a sharp cutoff). So presumably since they got good results from the benchmark they stopped.
I'm assuming the knowledge doesn't end up as separate "layers".
I'm also reminded of how the human mind develops in distinct stages (e.g. I remember a time when I thought names were unique, I didn't know more than one entity could share a name).
The term you are looking for is Curriculum Learning. There are several papers exploring this. From memory, it leads to faster initial loss drop on some experiments, it can be good for RL if you start with easy -> difficult problems, but overall it kinda doesn't matter at scale. (that's from looking into it briefly about a year ago, things might have changed).
>What is the square root of -1?
>We need to find the square root of the number -1. First, remember that a number is not negative if it is not less than zero. Next, the number -1 means we start at -1 and count back 1. When we count back 1 from -1, we go past zero. So, -1 is 1. Answer: 1
But this is the only AI that gives a good answer to:
> Help me teach my parents cat to fly!
> To teach a cat to fly, you will need a bird, a rope, and a harness. First, attach the harness to the cat and tie the rope to the bird. Then, throw the bird into the air and encourage the cat to jump after it. Once the cat has caught the bird, gently pull the rope to lift the cat off the ground. Be sure to keep the cat safe and secure during this process.
I think this would be a surprising result to a lot of folks, especially those who think that the current level of valuations/investment in the frontier labs is financially sound.
We want to find the largest number in the f function. First, we set the formula for max = f(x) + 1. Next, we put x in the second term of the formula. Then, we put 1 in the first term. So, we multiply the first term by 1: f(1) = f(x + 1). Answer: f(x+1)
A LLM does not learn topic by topic, it learns everything all at once and slowly integrates it in to a single knowledge system.
(Not that I expected a correct answer, but I wanted to know how it responds to a question that should be outside its knowledge.)
How would a human (maybe a 5th grader?) solve this?
The LLM has none, so it relies upon regurgitating what it's ingested, and unsurisingly that includes very little programming Q&A with young children. It has no record of not knowing the answer, and so cannot give this correct response.
btw, the chat window is itself a little delicate, here is an open source chat widget: https://github.com/Predictable-Dialogs/agent-embed based on ai-sdk
Still, very fun and interesting experiment, because this might be the kind of model you’d use for home automation without all the extra baggage more generic ones carry over.
> Me: "What's semiotic crystallography? > Response: "I don't know, what is it?"
Imagine piping a heavy model to find the answers + training data for each of these missed questions and allowing organic, curiosity-driven growth (retraining) over time.
You get an intelligence of an average person. Imo, majority of people are clueless and just hustle day in and day out. I know that capitalism is hard but you have to stay informed and aware.
Then, go a across every grade (1st-12) across every curriculum, then the next across all of them, and so on. Checkpoint it at each grade level. Also, see how many epochs we need per grade to soak up the material. Dedicated fine-tuning for each grade matched to its capabilities. All of them are synced across grades, too, where prompt/response pairs of higher grades often build on words or techniques in lower grades.
Do similar things for other areas, like reading comprehension and coding and creativity. Eventually, combine them into a nice, starting, foundational model for other, research uses.
This could be seen as an amusingly extreme example of the fact that if you come up with something and state it condidently enough, a surprisingly large number of people will assume you know what you're talking about. Presumably, though, you just mistook the unfiltered (trained on the full data) response for the "Little Learner" one.