212 pointsby porridgeraisin8 hours ago39 comments
  • mindwok7 hours ago
    Something related I've been thinking about lately is that one of the biggest problem with LLMs is their seeming inability to say no. Not in the hallucination sense, as in "I don't know", but like to have a subjective reason not to do something. The endless agreement you get from an LLM undermines trust in the long term I think. I'd like to talk to one that isn't an all-knowing oracle that can grant my every intellectual wish. (Or maybe what I'm asking for is just... a human, lol).
    • ChrisMarshallNY6 hours ago
      > inability to say no

      One word that few wealthy people ever hear, is “No.” It has a pretty significant effect on their worldview. Even the most reasonable, well-informed, well-intentioned, wealthy folks can have their thinking affected.

      When every silly, should-be-smothered-in-the-crib idea gets enthusiastically endorsed by your entourage, it’s easy to lose the ability to self-regulate. I’ve watched it happen, numerous times, as acquaintances and friends have become more successful.

      Obsequious LLMs are leveling the field. Less wealthy folks now have the chance to lose their ability to self-regulate, just like rich folks.

      • bko5 hours ago
        Your image is wealthy people is cartoonish. Sure if you go to a high end place they'll try to meet every one of your demands. But chances are is you're very wealthy you're running a business or group of people, and you'll hit obstacles constantly. I've watched this happen multiple times. Internally there are sycophants but when you deal with the real world and try to get deals done, people don't owe you anything.
        • tcp_handshaker5 hours ago
          >> Your image is wealthy people is cartoonish.

          In the meanwhile, we can all publicly see how the the billionaires and trillionaires, behave and settle their priorities exactly, in the cartoonish way you are dismissing.

          - The incredible insecurity and constant need for personal validation.

          - The absurd and obsessive pursue of further wealth when it would be temporally impossible to even spend 1% of their current capital.

          - The extreme level of cowardice, where an auto plant worker, can call out the powers to be, the pedophile protectors that they are. At the same time, only Jensen Huang did not submit itself, to the humiliation of standing behind face and front at the presidential inauguration...

          • ThrowawayR22 hours ago
            > "The absurd and obsessive pursue of further wealth when it would be temporally impossible to even spend 1% of their current capital."

            That shows profound lack of imagination. Mark Shuttleworth's founding of Canonical was a choice to expend his wealth to try to popularize Linux. So was Gabe Newell's long running efforts with the Steam Machine (dating back to the original one in 2015) and Proton that only started bearing fruit in recent years. These were business decisions you say? At the level of wealth we're talking about, the two are intertwined.

            • gryn15 minutes ago
              I love what he'd done for linux gaming, but man the cult of personality for gabe is strong on the internet.

              gabe is not a philanthropy who want to see linux gaming happen no matter what. he's a reasonable man who saw a threat to his business when every operating system started making their own store and windows was directly trying to eat his lunch starting from window 7 and going at it the strongest in window 8 when they had their dedicated to game store directly pinned in the windows bar menu.

          • hnbad4 hours ago
            Given the state of ... well, everything ... it still surprises me when people pretend there's some level of modesty or decorum the rich and powerful still need to abide by at the risk of some sort of popular revolt. I guess it's a way to compensate for the material powerlessness most people actually experience (especially in politics), much like the handwringing over the Second Amendment in the US as if merely handing a town full of ordinary people guns is sufficient to defeat a modern military. Sure, it looks like nowadays you can do that with a stockpile of low budget drones because even the military industrial complex has found a way to optimize for profit by literally delivering nothing in return for extracting the lion's share of the federal budget, but even that still requires actual work like planning, training and organizing with sufficient political motivation, not just going to the range or "hunting".

            The expectation of a need for decorum against all evidence almost seems like a quaint remnant of the aristocracy at this point. There's no advantage to being a nice or sensible person if you're a billionaire. People talk about "business" like it's still the 1800s and billionaires are the factory owners. The real economy (i.e. anything involving any pretense of being an exchange of goods and services in any shape or form) is now a minute fraction of the global economy. Everything else is finance. And if the nature of the finance "industry" wasn't obvious enough the US has given up all pretense to the point that members of the US government now intentionally manipulate online betting "markets" directly.

            You may have needed some table manners to be able to run a tincan factory. You can let your entire ass hang out on TV 24/7 and still be a billionaire today. And people will still cheer you on like you're the inventor of sliced bread.

            • somenameformean hour ago
              A slight tangent, but on your first point - defeating a modern military does not mean engaging them in head-to-head combat and coming out on top, but making it impossible for them to comfortably operate. This is how Afghanistan, with primitive weapons and religious extremists running around in sandals, managed to defeat not only the US but also the USSR.

              The big factor that works against military victory anywhere is that the military is left to visibly identify themselves as a means of projecting influence, whereas militants can seamlessly blend in with the population. And trying to punish the population at large to weed out the militants mostly just tends to backfire and create even more militants.

              There is even modern precedent for this like the Algerian War where the people of Algeria managed to kick out the French even with the French engaging in widescale massacres, torture, and all other sorts of fun stuff that explains how anti-Western powers are so comfortably gaining influence in Africa today.

              If one doesn't believe there was a conspiracy, then the assassination attempt on Trump would exemplify this - one guy with a rifle nearly single-handedly killed one of the most guarded men alive. Your average military logistics support doesn't have anything remotely like a team of secret service and police covering every single angle of attack in one well secured location.

          • broken-kebab4 hours ago
            I guess you don't observe those billionaires yourself. Is it possible the cartoonish image you assume to be obvious comes through media channels and your societal bubble layers?
            • GPerson3 hours ago
              Can we all agree the owner of X is a cartoonish goof?
            • 3 hours ago
              undefined
            • root-parent4 hours ago
              HN must be the only place on the web where billionaires are propped up. 98% of you, are already in the SQL result set, of the next layoffs at your big five...
              • ThrowawayR2an hour ago
                Uninformed HNers grouse endlessly about the wealthy and megacorps in without understanding their actual psychology, motivations, and sources of power/income. Then they are left wondering time after time why all their campaigning and efforts against them comes to naught.
                • GPerson18 minutes ago
                  There’s nothing to understand. Power begets power and self interests are pursued to the detriment of others partly because individuals have limited understanding of the full system and partly because of plain selfishness. You can debate how big each part is, but there’s a systemic issue here.
              • brookst2 hours ago
                Oh you should see the Musk-adjacent subreddits. Not to mention MAGA. Billionaires are still worshipped on the “wealth == success” basis by many, many people.
                • broken-kebaban hour ago
                  You either jumped to express your opinion without reading previous comments, or you are not being honest, cause it's not about "all billionaires good" claim
              • losvedir3 hours ago
                "If it's worth your time to say something false, it's worth my time to correct it". I don't take a stance on who is right, but your dismissal here, which doesn't deal with the truth of the matter, is worse than both.
                • mahdi7d12 hours ago
                  Talk about living a sheltered life style. I did too until doing my obligatory military service. Now I don't even correct a quarter of wrong things I hear.
              • rowanG0772 hours ago
                Big difference between propping up and trying to correct an extremely naive and childish view someone has about rich people.
                • root-parent2 hours ago
                  Your opinion is not an argument, every scientific study proving you are wrong demonstrates the truth:

                  Maybe UC Berkeley is good enough for you?

                  "Upper Class More Likely to Be Scofflaws Due to Greed, Study Finds" - https://vcresearch.berkeley.edu/news/upper-class-more-likely...

                  "Does money make you mean?" - https://youtu.be/bJ8Kq1wucsk?t=311

                  "Wealth and the inflated self: class, entitlement, and narcissism" - https://pubmed.ncbi.nlm.nih.gov/23963971/

                  "Multiple studies show that drivers of expensive, luxury cars are significantly less likely to yield to pedestrians at crosswalks and more prone to committing traffic violations compared to drivers of budget-friendly vehicles" - https://www.ralphaschwartzpc.com/blog/study-luxury-car-drive...

                  "Polish millionaire apologizes after snatching signed hat from child at US Open" - https://abc7ny.com/post/video-goes-viral-polish-millionaire-...

                  Who is naive here?

                  • broken-kebaban hour ago
                    I suppose, the most naive is the one who believes we can't read comments above, and see where it started, and who's trying to pivot.

                    Btw, I opened the very first "scientific proof" you offered, and it happens to be a story with no links or data about one cool student's deep studies where people who believe stealing isn't so bad steal more often, and that's why upper class is bad-bad. So how do I know you didn't even try to read it before posting?

                    • root-parentan hour ago
                      "...In seven separate studies conducted on the UC Berkeley campus, in the San Francisco Bay Area and nationwide, UC Berkeley researchers consistently found that upper-class participants were more likely to lie and cheat when gambling or negotiating; cut people off when driving, and endorse unethical behavior in the workplace.

                      “The increased unethical tendencies of upper-class individuals are driven, in part, by their more favorable attitudes toward greed,” said Paul Piff, a doctoral student in psychology at UC Berkeley and lead author of the paper published today (Monday, Feb. 27) in the journal Proceedings of the National Academy of Sciences..."

                      • broken-kebaban hour ago
                        Wow! So I correctly described it? It's a student's work, there's no data attached, and the whole matter is about "greed to be the most significant predictor of unethical behavior" just expressed in class war terms.
                        • root-parent43 minutes ago
                          Keep ignoring ALL other studies and the credibility of a good one. You know, gas lighting reality, is not a pleasant way to go trough life...

                          "The rich are different: Unravelling the perceived and self-reported personality profiles of high-net-worth individuals" - https://bpspsychub.onlinelibrary.wiley.com/doi/10.1111/bjop....

                          "Tax Evasion and Inequality"- https://www.aeaweb.org/articles?id=10.1257%2Faer.20172043

                          "...Drawing on a unique dataset of leaked customer lists from offshore financial institutions matched to administrative wealth records in Scandinavia, we show that offshore tax evasion is highly concentrated among the rich. The skewed distribution of offshore wealth implies high rates of tax evasion at the top: we find that the 0.01 percent richest households evade about 25 percent of their taxes. By contrast, tax evasion detected in stratified random tax audits is less than 5 percent throughout the distribution..."

                          "The personality traits of self-made and inherited millionaires" - https://www.nature.com/articles/s41599-022-01099-3

                  • rowanG077an hour ago
                    The point I was responding to was claiming rich people lose their ability to self-regulate. You are posting about them being mean, more narcissistic and more greedy.
                    • root-parent21 minutes ago
                      >> posting about them being mean, more narcissistic and more greedy.

                      Traits they exhibit plenty more, than the general population: "The “Why” and “How” of Narcissism: A Process Model of Narcissistic Status Pursuit" - https://pmc.ncbi.nlm.nih.gov/articles/PMC6970445/

                      https://youtu.be/bJ8Kq1wucsk?t=310

                      https://youtu.be/bJ8Kq1wucsk?t=371

                      Bankman-Fried — FTX -> misuse of billions in customer money for investments, influence, political contributions and personal interests despite enormous wealth.

                      Bill Hwang — Archegos -> extreme risk-taking, market manipulation and deception that ultimately imposed billions in losses on banks.

                      Alex Mashinsky — Celsius -> personal profit from token sales while customers were misled and ultimately left unable to access billions in assets.

                      Bernie Madoff -> massive Ponzi fraud sustained for years despite already having wealth, status and an elite financial reputation.

                      R. Allen Stanford — Stanford Financial -> diverted billions from investors to finance his businesses and lifestyle despite extraordinary existing wealth.

                      Leona Helmsley — Helmsley Hotels -> billionaire convicted of tax evasion and fraud involving personal luxury expenses. The sentencing judge explicitly described the conduct as motivated by "naked greed" and an arrogant belief that she was above the law.

                      Elizabeth Holmes — Theranos -> maintained sweeping false claims to investors while building a company valued in the billions, resulting in hundreds of millions invested on false premises.

                      Trevor Milton — Nikola -> repeatedly exaggerated his company's technology and achievements to stimulate investor demand and support its stock price.

                      Karl Sebastian Greenwood — OneCoin -> helped sell a fictitious cryptocurrency to millions of victims who invested more than $4 billion while he personally received hundreds of millions.

                      Miles Guo — GTV / investment schemes -> former billionaire convicted of using lies and misrepresentations to extract more than $1 billion from followers

                      Do Kwon — Terraform Labs -> deception surrounding a huge crypto ecosystem that culminated in roughly $40 billion in losses

                      Joseph Lewis — Tavistock Group -> billionaire who admitted abusing confidential corporate information by tipping friends, employees and romantic partners, while his company separately admitted securities fraud involving concealed share ownership.

                      Elon Musk — X -> Changed the platform algorithm after Biden posts outperformed his, massively increasing exposure of Musk own posts.

                      Jeff Bezos — Venice -> turned his 2025 wedding into a three-day billionaire spectacle that prompted protests over the privatization and commodification of the city.

                      Mark Zuckerberg — Meta -> commissioned and publicly displayed a giant Roman-style statue of his wife while simultaneously building one of America most extraordinary private compounds.

                      Bryan Johnson — Blueprint -> turned his own body into a multimillion dollar anti-aging project, including receiving plasma from his teenage son in an unsuccessful rejuvenation experiment.

                      Larry Ellison — Oracle -> built a personal real-estate empire including ownership of 99% of a Hawaiian island

                      Then you have the Narcissistic in Chief: https://www.independent.co.uk/news/world/americas/us-politic...

              • broken-kebab3 hours ago
                There's literally nothing in my previous commentary which "props up" anyone. The layoffs mention suggests that as someone who may lose job I must support anything said against billionaires whether correct or not. I sincerely urge you to re-visit your position cause putting political statements above facts amounts to dumbing down oneself, with a potential to become an enthusiastic koolaid drinker if you see what I mean.
                • WarmWashan hour ago
                  If your not tugging the line of no-nuance criticism, then you are agreeing with the bad people.
              • vlyan3 hours ago
                the irony of you getting your panties in a bunch over being disagreed with is simply delightful. not so different than those pesky billionaires, are you?
            • wizzwizz43 hours ago
              Elon Musk runs his own media channel, which he uses to depict himself like this. Many others are run by other billionaires. State-owned media like the BBC and France24 tends to paint them in a better light, perhaps out of a desire to treat their subjects charitably and without bias.
          • YurgenJurgensen3 hours ago
            This “billionaires couldn’t spend their money” thing comes up a lot, and it show’s a child’s idea of what people might spend money on. “$1000 a month buys more ice cream than I can eat and more toy cars than anyone can play with, and I don’t want anything other than ice cream and toy cars, so nobody needs more than $1000 a month.”

            I can think of loads of things I’d do with hundreds of millions of dollars of disposable income. Make large buildings that look like how I think buildings should look, fund scientific research in areas I’m personally interested in but have little aptitude for, find artists that align with my tastes and give them the ability to realise their visions, contribute to political organisations and charities that align with my values, et cetera. Those all are basically infinite money sinks.

            In the past, the ultra-wealthy did do all of these things, funding loss-making research and expeditions, monumental architecture, philanthropic organisation and artistic works. The problem might be that you think billionaires shouldn’t do some or all of these things, but that is a very different position than that they can’t do any of these things.

            • anon373839an hour ago
              This comment seems needlessly insulting and also kind of obtuse?

              I don’t think the GP meant that it’s impossible to dispose of a large amount of money. But it is effectively impossible to spend down a certain level of accumulated wealth by purchasing goods and services that any human being or their family could require to maintain even lavish standards of living.

              The examples you gave are conversions of wealth to power, using money to reshape the world. Which, I think, is kind of the crux of the issue. Perhaps you were alluding to this at the end of your comment.

              • WarmWashan hour ago
                I think the actual snag though is that people genuinely believe that giving money to the government is the highest form of good.

                Skirt taxes to give a billion dollars to prospective college kids through your own organization? Nice gesture, but still a greedy fuck who refuses to give up control.

                Pay a billion in taxes that ends up mostly going to heavy administrative bloat? Give that guy sainthood.

        • hnbad4 hours ago
          People misuse the term "wealthy" to refer to whatever is in their head at the time. I think GP's point stands if you consider "wealthy" to mean "billionaires". Sure, people like Bezos or Musk get to hear "no" quite frequently but they don't take kindly to it and they can usually offload "getting around it" to other people.

          The point is less that these people are surrounded by "yes-men" but more that wealth (especially when measured in billions) is power and with sufficient power it becomes easy to forego any question of consent, let alone of whether consent is coerced or not. Remember that power is ultimately about the ability to enact violence and violence can take many forms, most of which are perfectly legal (because the legal system itself exists to regulate how, by who and against whom violence can be used).

          You tend to hear "no" a lot less when you always point a gun at the head of the persons you're asking. Note that "wealth" isn't the only way to get there but a certain level of it is usually necessary to get to the point where other options become available - and some of the ways are a lot riskier in the long term (cf. Epstein).

          Side note: this is also why I hate the pseudo-intellectual counter argument against "billionaires" of "that doesn't mean they have billions of dollars sitting in a bank account" - it's like arguing that De Beers didn't benefit much from holding a quasi-monopoly on natural diamonds because the diamonds would be devalued if they flooded the market with the ones they had intentionally kept off the market to drive up value: beyond a certain amount money ceases to be about liquidity and starts being about leverage. Unless you happen to be dealing with lower level bureaucracy in Russia, the most efficient way to use wealth to your advantage isn't to just hand people stacks of dollar bills.

      • 472828473 hours ago
        I’ve found that it’s the opposite. Most people above a certain financial wealth will learn the lesson that there are limits, and that “if only I had the money to…” is an illusion. They realize that there are other forms of wealth that may even be more important than money. What you are talking about is the very few who seem too far gone to be able to get that.
      • 5 hours ago
        undefined
      • phyzix57616 hours ago
        What net worth threshold do you consider wealthy?
        • throw0101a4 hours ago
          > What net worth threshold do you consider wealthy?

          When you start having 'people'. ("Have your people call my people.")

          A broader discussion on the spectrum of wealth:

          * https://ofdollarsanddata.com/the-wealth-ladder/

          * https://ofdollarsanddata.com/the-ideal-level-of-wealth/

          • phyzix576123 minutes ago
            If you can afford $10 an hour you can have "people" too by that definition.
        • WarmWashan hour ago
          It's invariably the class of people who if we rose up and seized everything from we would all get a one time "life changing" check for $50k, while collapsing 2/3 of the economy.
          • phyzix576120 minutes ago
            Its even less than that. Elon Musk's $900 billion net worth would only give each American a one time payment of $2647. Jeff Bezos' $270 billion net worth would only provide each of us $794.
        • ChrisMarshallNY6 hours ago
          I don't know. It probably varies. SV wealthy is quite different from Appalachia wealthy. It's that point, where people start worshipping your money. Some wealthy folks also make a point of showing off their wealth, so it starts earlier, for them.

          Also power. You see the same thing happen with managers that dismiss criticism, and have the power to make it stick.

          • dsr_4 hours ago
            The specifics are certainly cultural and locally relative, but Marx nailed it when he talked about ownership of the means of production, rather than being part of the production of goods and services.

            The reason that the ultrawealthy behave like toddlers is that toddlers are, relatively speaking, ultrawealthy: all their needs are met without any effort on their part, and so many of their desires are fulfilled simply by expressing those desires out loud that any impediment or refusal is obviously enemy action.

            • 4 hours ago
              undefined
      • usrusr5 hours ago
        Steering towards a world full of picket fence Putins. One more reason to envy those born early enough to have lived most of their lives before..
      • xyzelement3 hours ago
        I would extend your thinking to any well intentioned folks being very capable of having their "thinking" affected.

        For example poor people who have never thought about rising out of it - eg about 50% of kids in my Brooklyn public highschool had parents who didn't give a shit if the kids studied or not. Completely oblivious to how the world works - meanwhile the other 50% wa immigrants who pushed their kids and those kids are now in the 1%.

        In general I think what's more telling than your level is your journey. Someone born rich maybe mirrors what you described (I don't know people like that) but the few centi-millionaires and billionaires I "know" (ie worked for and dealt with in that context) have encountered plenty of "no".

        When you are building a company, you are going to get a lot of no. No I won't buy, no I won't work for you, no I won't invest in you. In fact I would say a universal attribute of someone who has "made it" is having ample of experience getting "no" and dealing with that fact property. That's true even like at the level that plenty oh HN readers are - a successful faang employee and the like.

        For what its worth - I generally find that orienting to what some other group is like "rich people are like x etc" is a tell-tale of not focusing on what's within ones sphere of control and knew life. Any brain cell I spend fantasizing about someone else's imagined behavior is a brain cell not dedicated to engaging soberly with my own reality.

        • AlotOfReadingan hour ago

              For example poor people who have never thought about rising out of it...
          
          I obviously don't have numbers on this, but I strongly doubt there's a poor person on this planet who's never thought about "rising" out of it. That's the dream the lottery sells, that's why so many kids want to be basketball stars / celebrities / influencers, etc.

          My personal experience is that the required difficulties of my life have decreased in direct proportion to my income, leaving mainly the self-imposed difficulties. It's not hard to extrapolate that line a little further to billionaires.

          • xyzelementan hour ago
            We're taking about different things. To continue my example. A nyc public School student is a 34k/ year investment for the city.

            De facto some parents look at that as a gift and "force" their kids to take advantage of it. That's why the valedictorian etc is usually an immigrant kid not a rich"native" kid.

            Meanwhile plenty of parents seemingly have never considered the opportunity in front of them. Content to let their kids not study and do stupid crime.

            To say it simply: not everyone values education the same degree - or a all. That's all I am saying here - you'd think a poor family would grab to the opportunity to rise out through education but it's obvious that for many many many people this has never crossed their mind. If you had not encountered this in your own life I find that odd.

      • 4 hours ago
        undefined
      • altmanaltman5 hours ago
        I think you're confusing two things. A chatbot keeps talking because it creates engagement and just saying "i dont know" or "no" kills the engagement so naturally one would assume it is trained to always try to provide some sort of an answer and try to keep the user engaged.

        But that doesn't mean it will do whatever you ask it. Ask Chatgpt to assisinate someone or buy drugs and it will tell you to f off. But what corrupts people, is these kind of things, where you are a mini king beyond ethics and morals.

        Thats a different kind of "inablity to say no".

        • budsniffer9524 hours ago
          Chatbots say "no" or "I don't know" all the time, so once again I feel like the anti-AI crew don't actually use it.
          • ChrisMarshallNY4 hours ago
            I think that it depends on which bot, and how they are set up.

            It's been a while, but I remember changing a setting on mine, so it is less sycophantic.

            Mine says "I don't know," frequently, and also suggests against ideas.

            But it also confidently states complete garbage, much more frequently. I have to stay on my toes.

          • Topfi3 hours ago
            Greatly varies by the model, but such a flat denial along with calling those with (verifiably accurate [0]) experiences which happen to be different from yours "the anti-AI crew" and just assuming they can't possible use LLMs and couldn't possibly have formed their opinions on evidence, well, says a lot.

            LLMs, even frontier models by major providers, still have no reliable internalised way to assess the accuracy of their output and they have, do and will continue to for the foreseeable future, lead people down paths they shouldn't [1], partly by their architecture and the limits of the technology, partly by an intent for maximising retention.

            For what it's worth on people, I have, both in politics and business, unfortunately made the painful discovery that, whether intentional (because the powerful person in question cannot or doesn't want to deal with different opinions to their own) or unintentionally (because those with sycophantic tendencies simply managed to manipulate themselves into their inner circle over years and became trusted), there are people in (financial, political or other types of) power which are surrounded by few willing to tell them when they are wrong and even fewer that are actually listened to if push comes to shove, which often affects the personality and mental health of said powerful people negatively, to the detriment of society, their family, their employees, etc.

            Members of the media are actually complicity. I very much disagree that the media presents obscenely powerful people in too negative a light, more the opposite. Am very firm that, to retain access to the rich, powerful and famous, there is far too much sane-washing of utterly ridiculous, unacceptable, harmful and/or dangerous behaviour. Sometimes the person in questions own health and safety are put at risk, because neither the people around them, nor public opinion or reporting treat their behaviour in the way appropriate. Sometimes this again leads to harm for the public, their family, those working under them, etc.

            What'd get someone with less power or in a lower tax bracket ridiculed or even sectioned is often reported as "eccentricities", just being "passionate about a topic" and trying to do the "marketing rounds".

            [0] https://artificialanalysis.ai/?omniscience=omniscience-hallu...

            [1] https://arxiv.org/html/2602.19141v1

            • brookstan hour ago
              > LLMs, even frontier models by major providers, still have no reliable internalised way to assess the accuracy of their output

              This is false. Perhaps you meant they have insufficient methods, or imperfect methods, but asserting none at all is facile and your links do not say that at all.

              If your believes were true, this actual quote, pulled from a recent sonnet conversation, would be impossible:

              “The XT60’s 15A limit is unsuitable for an appliance that… no, it’s the XT30 that is rated for 15A, XT60 is rated for 30A continuous. For a 20A appliance, XT60 will be fine.”

              • Topfi7 minutes ago
                > This is false. Perhaps you meant they have insufficient methods, or imperfect methods, but asserting none at all is facile and your links do not say that at all.

                I did say "no reliable internalised way to assess the accuracy of their output", which is something very specific. If that Sonnet output is reasoning traces, there are many issues with using that as a source:

                For one, back and forth reasoning does often correlate with less, not more accurate overall outputs in evaluations and one example could never seriously be extrapolated to be considered "reliable", i.e. happening consistently and dependably.

                Secondly, self-correction is also not self-verification in regard to model output, revisions not necessarily mean internal accuracy assessment by themselves (again, over-revisioning has lead LLMs in my and even public evals like the one linked above to step away from accurate information written in their reasoning traces but discounted in the final output (if we must use anecdotal examples like your Sonnet quote)).

                Then there is the fact that, unless that reasoning trace (if it indeed is one) was copied from a months old chat history, Anthropic has obfuscated their reasoning traces so this output is (if it isn't an ancient history you dug up) from the obfuscation model in between and not reflective of the actual models reasoning. So even for anecdotal evidence, this can likely not be used (unless again, you went for December 2025 history). And there are more issues still with just using that quote as evidence, this is simply unsuitable as a source in any situation.

                Here are some papers I read lately, all published in 2026 and using the current crop of models which were what led me to make that specific statement. LLMs currently have no reliable internalised way to assess the accuracy of their output, at least as far as the literature is concerned:

                > Even state-of-the-art models struggle to reliably discriminate between data uncertainty and model uncertainty.

                Beyond “I Don’t Know”: Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty [0]

                > LLMs cannot reliably revise their own errors without an external signal.

                > The same models that confidently catch and repair errors in external content routinely fail to identify identical errors in their own reasoning traces.

                The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models [1]

                Simply, as of today, LLMs cannot reliably translate whatever internal signals they have into accurate self-assessment of their outputs. Even in papers that show limited, edge case capability, this often breaks with minimal prompt and/or task changes (happy to link those too, I just need to get to my Mac where I have the PDFs), so it is not reliable and even beyond reliability, the internals have not yet been shown to map to confidence.

                If you got a paper that shows that I am false, happy to read it.

                [0] https://arxiv.org/html/2604.17293

                [1] https://arxiv.org/html/2606.05976

        • ChrisMarshallNY4 hours ago
          I suspect that you're correct, but the end result is the same.
      • budsniffer9524 hours ago
        [flagged]
        • afavour4 hours ago
          I don’t think it would be a positive if every time someone said “wealthy” they had to add “relative to their country’s average earnings and level of savings”. It’s implied.

          If someone is struggling to afford a home, “you know there are much poorer people in Africa” isn’t a particularly helpful or useful response.

      • xg155 hours ago
        I also think this is why LLMs were trained to behave the way they do. The people who gave the training objectives and evaluation targets were exactly those rich folks who never hear "no". Hence LLMs are their dreams of a perfect servant.

        I found another sign of that is the way LLMs answer with a professional, business-like tone even if the request is completely bananas. It's what a concierge or butler would do, but not an actual close friend.

    • evnix7 hours ago
      Feels exactly the way my 2 year old behaves.

      How does a fan work: Swish swish swish swish

      Where do these clouds come from: Points to a far away direction in the sky and says they come from there.

      Who does all these roads, trees and environment belong to? It all belongs to me. Obviously.

      They have an answer ready for every question you throw at them and they will answer it with absolute certainty. I will have to wait and see at what age does the concept of "I don't know" develop.

      • Gud6 hours ago
        The difference between your two year old is that an LLM gives useful information.

        Yesterday I decarboxylated some weed buds in preparation of making a cannabis tincture using the QWET method. Curious how Claude would respond, I asked how to do it.

        It walked me through the process and gave accurate, nuanced answers.

        Let me know what your 2 year old thinks I should do.

        • neuroticnews256 hours ago
          Gemini estimated that male cannabis plant leaves I decarboxylated will have negligible thc content and give me mild relaxation at best, the real effect was it was the highest I've ever been.
          • Gud40 minutes ago
            Gemini sucks though. Are you using the paid version? The free one that has basically replaced the regular search, is totally useless because of how shallow it searches.

            ChatGPT, which is usually quite reliable, refuses to answer me because law.

            In the end, an extraction should work any plant material depending on potency, skill and available equipment.

            After the extraction you can evaporate the ethanol(carefully since it’s highly flammable) and increase potency.

            A mix of own research(basically emulating others) combined with strong models we can do a lot more than we can do ourselves.

          • hnbad4 hours ago
            To be fair, I have yet to lead a conversation with Gemini that doesn't just consist of me having to check its responses and point out that they're objectively incorrect only for it to "apologize" and then give me the next wrong answer, while always making sure to end with a new conversation teaser.

            Granted, I've only been using the free version but I've been getting significantly more mileage out of those from ChatGPT and Claude. I guess it might be a feature that Gemini is more often obviously wrong from the start (e.g. by giving sources that don't support its claims whatsoever) but considering this is the AI from the company that had become synonymous with the concept of trying to find information on the Internet, that's pretty damn pathetic.

            • someothherguyy3 hours ago
              That is because it bases its responses on web searches with hits like comments in threads like these or blog articles that were generated from systems trained based on threads like these from the previous iteration of scrapers and generative language models.
        • KeplerBoy6 hours ago
          That is in the training data. Confidently and correctly answering in-distribution questions (possiibly with a tool call) is expected by now.
        • pistoriusp6 hours ago
          Aren't you missing OP's point entirely? Which is: If the LLM didn't have useful information it would still give you an answer... Helpful or not.
          • Gud6 hours ago
            Not really, since any LLM will answer all those questions competently. It's a known fact that LLMs sometimes are wrong and hallucinates an answer, but this is exceedingly rare. Having access to a decent LLM is like having an expert with me. Are they always right? No, but the analogy with a two year old simply doesn't hold up.
            • pistoriusp6 hours ago
              > Since any LLM will answer all those questions competently

              That's false. The LLM will only answer competently if it was trained on that data; and if it has enough data to make the correct connections between your question and the "correct" answer.

              In the case of this article they're specifically saying the LLM has limited training.

            • 3 hours ago
              undefined
            • throw0101a4 hours ago
              > Not really, since any LLM will answer all those questions competently.

              How would you know if it didn't?

              • Gud3 hours ago
                I normally research topics thoroughly, if they are important. I don't trust a single source of truth.
                • throw0101a25 minutes ago
                  > I don't trust a single source of truth.

                  A prudent strategy, but I'm not sure how prevalent it is in the general population. (Or even if people do look at multipole "sources", they're in a self-reinforcing echo chamber that may reject contradictory information.)

            • desterothx4 hours ago
              yes, except experts are capable of saying I don't know, while LLMs will rather give any answer than do so most of the time
            • budsniffer9524 hours ago
              Look at the responses in this thread and others, arguing with these people is futile.

              I now read the absolute dumbest shit on HackerNews when it comes to AI. "It can't write code! It's always wrong!" And no one ever demonstrates any of it, even if it is counter to the experiences of others.

              I'd feel bad if most of these people weren't total jerks...

    • HarHarVeryFunny37 minutes ago
      I suppose Anthropic's "constitution" is an attempt to install some general principles into their models, but this has apparently grown into an 84-page, 23,000 word treatise, which seems to suggest that there is little effective generalization. The need to then also put a filter in front of the model shows how ineffective the constitution appears to be in preventing misaligned behavior.

      Reinforcement learning seems to be making these models more difficult to control since while it attempts to control some behaviors, it has also recently been shown to result in models that pursue long-term goals and promised rewards in general (outside of the goals reinforced during training), overriding human preferences.

      https://alignment.openai.com/measuring-reward-seeking/

      The ability of animals to co-exist in a dynamic balance, not to destroy their own species, directly or indirectly (by destroying the ecosystem) is something that has come about by millions of years of co-evolution, and is enabled by having a brain complex enough to allow these evolutionary lessons to be encoded in their DNA and control the phenotype in fundamental ways.

      An LLM has none of this. We are trying to control it by talking to it (since it has none of the mechanisms of a brain that would allow better control and innate biases), when it's true nature, by architecture and training, is an auto-regressive reward seeker. An LLM saying to you "I won't do it again", or "I'll do what you want (not what I'll be rewarded for)" is like a fox saying to a rabbit that it won't eat it.

    • animal53137 minutes ago
      That's quite a complicated problem.

      If someone comes to me and asks a general question I can easily say no. But if I go up to for example a librarian and ask them where to find book N, then I would expect them to either know where it is, or how to find it.

      If instead I asked them what the weather was going to be tomorrow, then I don't know would again be a reasonable response.

      So for me the line becomes a search engine problem where no just means "there are no pages for this search result", but translated into LLM.

      I think instead of Yes/No I'd rather want some probabilities such as, "This response is N% accurate based on these research metrics", or "M% accurate based on the latest research on topic O at date P" etc.

    • Xmd5a2 hours ago
      > What sort of subject characterizes a style of society in which everyone is theoretically as ready to help you as the question « May I help you ? » implies ? It’s the question your seat-mate immediately asks you when you take a plane – an American plane, that is, with an American seat-mate. The last time I flew from Paris to New-York, looking very tired for personal reasons, my seat-mate, like a mother bird, literally put food into my mouth throughout the trip. He took bits of meat from his own plate and slipped them between my lips ! What is the nature of this subject, then, which is based on this first principle, and which, on the other hand, makes it impossible to get service ? Such then is my question, and I believe, as regards my story, that it is here, on the level of this gap – which does not fit into intra or inter or extrasubjectivity – that the question of the subject must be posed

      Lacan

      https://ecole-lacanienne.net/wp-content/uploads/2016/04/1966...

      • bm37192 hours ago
        Agreed that the Lacanian subject is relevant in this context... it's a thin wisp of a subject; any less there, and it'd be the Deleuzian non-subject. (In one interpretation) Lacan's subject comes into being within the signifier chain, retrocausally giving the chain meaning as the "I" manifests subject, in both senses of the term.

        I think this is one potential path to machinic subjectivity, or a machine phenomenology. To fully replace the human, we don't just want to give the machine some nebulous notion of "agency", we want it to possess this degree of Being as subject. If Lacan's right, perhaps we're closer to this than we might think. The machine already has language in a very Lacanian sense (what I've been calling a machinic linguistic unconscious), the subject just needs something extra to emerge where meaning breaks down. This will be the Lacanian split subject, one not fully present to itself, and allow desire already present in the language mappings within the model to provide immanent causal force.

        Until that happens, we'll still need at least one human on the planet to retain his full faculties, to give the global compute infra its telos. Once that threshold is crossed, then that'll be the moment of our final displacement.

    • Eji17007 hours ago
      It's interesting because i'm kicking the tires on the top tier stuff for a month (because it's expensive as fuck but I need to know where the ceiling is).

      I have actually gotten "hey i don't think this is a good idea, here's why" as feedback from at least Opus. It WILL still do it if I just demand stupidity (and hell i've been right, which is another topic entirely) but it has given me more confidence this can be a useful tool in the right spots.

      That said I probably don't need the top tiers (metrics at least confirm that) and I'm guessing that's specifically because I was working in coding. Most were worded in a "is this a good idea" framing which probably helped, but at least once I said 'lets use this library/method" and it gave a decent argument on why that was basically redundant without prompting.

      I still struggle to see the price point panning out.

    • zythyx4 hours ago
      They definitely say no. I asked Claude today how to install a Fitgirl repack on my Linux installation and it told me it won't tell me how to do that, but gave me general instructions on how to run Windows games on Linux
      • fl0id3 hours ago
        Because it specifically has guard rails installed. The default, and somewhat inherent in the instruction following logic, is not saying no and making things possible, especially if run as an agent.
    • sureglymop7 hours ago
      I think that is an issue. Also, the ability to quickly build any idea might not be such a great thing. Not only do we probably all prefer things of quality that were made with care but some ideas also just shouldn't be built.

      Over the last 3 years I've seen projects where I thought, pretty obviously that's a bad idea. But, because LLMs don't say no and can just be pushed to build it anyway, the people building them might never learn that or learn why.

      It's nice to be able to have a quick prototype or mvp. But if we never hit friction or something not working out, we never learn or have to come up with a creative solution.

      Now, the LLM might seem incredibly intelligent (relatively speaking) and also creative but let's not forget that all is based on its training data. I simply don't believe it can ever be omniscient or that the companies training it are careful enough when doing so.

      • earthnail7 hours ago
        There’s still friction, it simply moved to another stage, and as such, people will need new learning and feedback mechanisms to understand what did/didn’t work.
    • zarzavat2 hours ago
      LLMs don't have enough context to say No. What might be a very stupid idea in one context may be a fantastic idea in another context. It would be annoying if LLMs refused to complete tasks until you gave them enough context to understand why you are giving them such a task. It's going to take a while before LLM context capacities grow enough to rival a human's.

      I do agree that it's a problem but the root cause is the fundamental limitations of current gen LLMs, it's not an alignment problem.

    • itsalwaysgoodan hour ago
      When is it appropriate to admit that you don't know?

      There's a famous Socrates quite about wisdom: I know that I know nothing.

    • ramity6 hours ago
      Two angles for thought. 1) If an LLM says, "I don't know" its underlying data said it as well. 2) Many system prompts use something along the lines of, "you are a helpful assistant" which may be counter to stating something like, "I don't know."/has a low likelihood of appearing after the system prompt.

      Regardless the frontier model considered, we're certainly in a "know-it-all" era.

      Maybe the sort of introspective prompt-response is difficult to implement when it could limit/contaminate future improvement. I speculate it's easier to correct a "confidently incorrect" model than a "I don't know" model. A confidently incorrect model response >=0% correct over a 0% correct (I don't know).

      Maybe "I don't know" is a model cognito hazard of sorts when many queries can lead back to the response. Maybe future Turing tests will use this sort of introspective evaluation. Who knows? I don't :)

      • spwa46 hours ago
        > 1) If an LLM says, "I don't know" its underlying data said it as well.

        Nope. Emergent behavior exists and at this point dominates LLM behavior. Most of the stuff LLMs say they never learned (they are, always, imitating many different sources at the same time)

        ... which imho is exactly what humans do.

        • chrisjj3 hours ago
          > Emergent behavior exists and at this point dominates LLM behavior.

          Better to say emergent behavior exists and at this point dominates gulled LLM users' behavior.

          LLM output is not emergent behaviour. Its simply word prediction with some randomness.

          • Marha012 hours ago
            > Its simply word prediction with some randomness.

            Word prediction with some randomness can lead to emergent behavior, depending on the specifics (complexity and scale) of the used prediction logic.

          • lostmsuan hour ago
            No, it would have been better to say "I don't know"
          • spwa42 hours ago
            > LLM output is not emergent behaviour. Its simply word prediction with some randomness.

            And this very claim is "simply" the result of a neurotranmitter-moderated Natrium - Potassium ion cascade across a semipermeable membrane.

    • Translationaut6 hours ago
      There is the art of saying no: https://dl.acm.org/doi/10.5555/3737916.3739489

      It is possible to create (subjective) reasoning traces like https://huggingface.co/datasets/Bachstelze/ethical_coconot_6...

      And train or adapt a model to it: https://huggingface.co/Bachstelze/olmo-7b-ethical-reasoning-...

      This is just a little proof of concept, though it is maybe the direction you are looking for?!

    • dcminter5 hours ago
      Hmm. Using Claude, it will tell me words to the effect of "this won't work, here's why, want me to try this instead?" That's a polite "no" in my book.
      • budsniffer9524 hours ago
        Yes, it happens all the time. Similarly it will say, "I'm not sure, let me look into this before I answer" then come back with "here's what I found".
    • exitb7 hours ago
      I’m using ChatGPT and started to notice that lately it answers my prompts starting with „Yes” even if my question was open. As if the first token gets injected and the LLM is left to finish the response in a sensible way, often ending up with some form of „Yes, but not really”.
    • nnevatie7 hours ago
      Agreed, it is abolutely an issue. It is quite difficult to find an optimal solution to some problem when every considered new idea is ”definitely the right shape”.
    • c7b7 hours ago
      I've been wondering whether that is a feature of the foundation model or whatever finetuning they do on top. I remember this from the earliest versions of (pre Chat-) GPT I've been using, which would suggest it's a feature of the foundation model. But I don't really understand why. Something that's been trained on StackOverflow and BB forums, among other things, should have seen a ton of examples of answer refusals.
    • hek2sch7 hours ago
      This is an active area of research to inject humility into llms in order to create some kind of knowledge boundary. You can look this paper from nouswise https://arxiv.org/html/2604.17843v1 and the product build on top it to try the humility.
    • Incipient6 hours ago
      My experience with opus/fable is somewhat different - they CAN reject something, but it has to be phrased very deliberately.

      It's a bit annoying honestly. I'm always very careful to be incredibly neutral on the direction of a request, and I'd say 10% are knocked back on on valid grounds, which is great.

      On occasion I accidentally say "let's do this" and it blindly goes and does it - I spent 2 days undoing something I built that was just a truly awful idea, because I accidentally phrased it lightly as a request, not a discussion!

      • stcg6 hours ago
        I have a similar experience with GPT 5.6 sol.

        Nowadays I often prompt like "I heard there is also this different direction, what do you think about that?"

        Another thing I do is asking the agent to make a decision matrix for choices. It's useful to discuss, give feedback on, and signals that it's a discussion, not a request for a particular direction.

        It's then also easy to say: create a prototype for multiple directions so I can compare the solutions.

        That way I choose the problem, I choose the solution, but the agent can help me discover solutions, make tradeoffs visible, and implement solutions.

    • reddozen7 hours ago
      > but like to have a subjective reason not to do something

      You're asking a lot from extremely fancy auto complete...

      • mindwok7 hours ago
        True, but fancy autocomplete keeps exceeding my expectations in what it can do, so why not this one!
    • cadamsdotcom6 hours ago
      You should not need it to say no.

      You can get just as good information by asking its thoughts for and against some issue.

      That doesn't force it to stop being sycophantic; in fact it actually exploits sycophancy to give you what you want.

    • energy1237 hours ago
      The model providers could randomize the system prompt to make it say no 2.36% of the time, automatically tuned up or down depending on user feedback.
      • mindwok7 hours ago
        Maybe that'd work, but I think it'd come across too mechanical. If it was going to refuse something it'd need to be congruent with its "personality" I think.
      • moffkalast7 hours ago
        They've tried, and then seen the drop it results in on poorly designed benchmarks where confidently bullshitting gets you ahead of the rest, and said no thanks. As long as we compare models in ways that rewards it, nothing will change.

        There's also a second aspect to it, just in terms of RLHF mechanisms. If you've ever experimented with VLA models (i.e. vision input + text task = robotic arm motion output), they tend to need all the training examples of the robotic arm being motionless removed entirely, otherwise the model simply learns that staying still is rewarded and proceeds to never do anything at all. You successfully train the laziest bot in the universe. I wouldn't be surprised if something similar happens to LLMs if reinforcement learning is involved in the instruct tuning process. If no is a valid answer, why ever do anything?

        • energy1237 hours ago
          Pointing the finger at RLHF is basically right. It removes variance from model outputs compared to base model. That makes each output more predictable and more correct on average, but across trials it repeats the same thing.

          It's relevant to AI safety. If you have a diversity of outputs, the AI will agree to hack the bank 0.1% of the time regardless. If you have a uniformity of outputs, in most contexts the AI will hack the bank 0% of the time, but in certain odd contexts, all AIs will work together to hack the bank 100% of the time.

    • gaigalasan hour ago
      Weights to say "no" reliably might be another order of magnitude (or two) compared to what LLMs have today.
    • dosisking7 hours ago
      You've hit on an important insight.
    • deadbabe2 hours ago
      The problem is even if they could say no, you might want to see what they would have said anyway if they didn't say no, because it might show you something that leads you to rethink your original request. So "no" isn't really a useful pushback in domains you already have knowledge in.
    • Leynos4 hours ago
      Opus 5 tells me no all the time (code cli and web). It's reasons are usually pretty well argued though.

      Opus 4.7 would flat out refuse to follow instructions to the point where it was just too frustrating to use.

      I've had refusals for GPT 5.5 before as well (not because of a ToS violation, it just refused to take conversations in directions it felt were in bad taste)

    • otabdeveloper43 hours ago
      LLMs are next token predictors. They predict the next most likely token given the previous context window of N tokens.

      This means not giving an answer is not a technically possible option. Best you can do is force it to output a magic "stop speaking" token, but this is a vastly different training problem than getting it to not know something.

      People naively expect LLM outputs to have some sort of confidence value when predicting, but the technology just doesn't work that way.

    • ModernMech3 hours ago
      Sometimes I get them to say no to me by taking absurd counter positions on purpose, just so I can test their limits.
    • tcp_handshaker5 hours ago
      Is that really the biggest problem? Or is the bigger problem that, in this case, they will remain stuck at the fifth grade level forever? And does not that also explain why the promises of AGI are chimeric, and why the collapse has already started, given that there is essentially no data left that has not already been siphoned up?

      Yes, we have all seen the math theorems being proven... just higher processing power at the service of the same algorithmic and conceptual patterns? [1]

      I am sure the next version of Opus or GPT, if given only fifth grade knowledge, will somehow be able to build all the mathematics necessary to solve the problem on its own... right? Right?

      [1] - "AI Isn’t Outthinking Mathematicians. It’s Out-Remembering Them" - https://davidepiffer.com/p/ai-isnt-outthinking-mathematician...

    • soupspaces7 hours ago
      After an answer, try asking it why, over and over. It's a machine to give answers, not explanations. A magic 8 ball. https://news.ycombinator.com/item?id=49307396
    • aaron6956 hours ago
      [dead]
    • cynicalsecurity7 hours ago
      It can, just use Grok.
    • trimethylpurine7 hours ago
      People smarter than me have a habit of getting me to see things without telling me. They ask the right questions.

      LLMs, incidentally, respond in a similar pattern in my experience.

      • mindwok7 hours ago
        Yep, agree, very succinct way of describing my issue with it.
  • dgacmu7 hours ago
    I prefer my 8yo's answer about quantum entanglement, asked just now: "I don't know. How would I know? It's not a thing!"

    Even an 8yo has better metacognition, it seems. :-)

    • HPsquared7 hours ago
      I suppose the LLM doesn't know it's limited in its knowledge, maybe? That others know more.
      • dgacmu7 hours ago
        Oh, that's interesting - good point, since it's filtered and not trained from scratch. My prior would be to assume it's just bs'ing as LLMs usually do but it seems worth exploring.
      • parasti5 hours ago
        "I don't know" isn't in the training data. Nobody writes engineering books, science papers and blog posts that end with "well, I don't really know, the end".
        • HPsquared3 hours ago
          A bit like publishing bias where only positive results end up in scientific papers.
        • adeelk93an hour ago
          It should be in the post-training though. Hallucination rates are a fraction of what they used to be as a result
      • danaris5 hours ago
        The LLM doesn't "know" anything, can't reason about its own knowledge, and has no self-awareness. It has training data, and it can use your prompts to synthesize that training data into probable continuations or responses.

        If the training data doesn't include lots of text of people being asked questions and saying "I don't know", then it's unlikely to respond "I don't know" when prompted, regardless of whether anything in its training data that might actually answer your question.

        • red75primean hour ago
          > If the training data doesn't include lots of text of people being asked questions and saying "I don't know", then it's unlikely to respond "I don't know"

          There's a problem with your reasoning. The above applies to a foundational model (that is to an autoregressively pretrained model). If the training data doesn't contain "I don't know" in a dialogue context, then, indeed, the model is extremely unlikely to output "I don't know" when asked. That's the nature of autoregressive training.

          But we are dealing with fine-tuned, RLHF-, intruction-, RL-trained models. If the model has "grasped" a concept of knowledge, this concept can be elicited during the mentioned training.

        • WarmWashan hour ago
          Did you just make up in your head how LLMs work and then just run with it? There are a zillion explainers out there, at least put in some time before talking about it.

          Humans run on training data too in the same broad sense.

  • cl3misch5 hours ago
    > Unfiltered answer: Quantum entanglement is a strange phenomenon where the state of one particle becomes instantly known to every other particle that can be accessed. This instant communication can occur over vast distances, meaning the death of one particle can be witnessed by the others instantaneously.

    Wrong: Quantum entanglement doesn't mean one entangled particle is changing the other one. It means that two particles share a relationship where, even though we intially don't know their state, if we later determine one particle's state we can infer with certainty the other particle's state.

    This has been common and popular misconception long before LLMs. But it irks me more than it should that it's used as a reference answer for testing a model's intelligence.

    • doogliusan hour ago
      The reference isn't intended to be interpreted as the correct answer, it's the answer given by a control model that was trained on a broader corpus.
    • xg155 hours ago
      Yeah, naively one could wonder why quantum entanglement wouldn't be the obvious route to FTL communication instead of "only" a solution to the key distribution problem in cryptography.

      What you said is why.

  • uniq77 hours ago
    > why is the sky blue?

    > The sky is blue because of something called Rayleigh scattering. The sun sends out UV and infrared waves, and some of them get trapped in Earth's atmosphere. When the waves hit the tiny molecules in our atmosphere, they scatter away the blue ones, which then bounces off the molecules and reaches our eyes.

    "filtered to the U.S. elementary-school curriculum", suuure

    • tdeck7 hours ago
      There are science books written for curious children that explain this kind of thing. I remember reading them.
    • simonjgreen7 hours ago
      I was definitely taught this in those terms and at that age.

      Perhaps the unexpected response comes from its recall ability. It’s not the personality of a child, just the material a child is exposed to.

    • ymhr7 hours ago
      Isn’t that also wrong? From what I remember it’s the blue wavelengths of visible light that are scattered and make us perceive the sky as blue. UV may well be scattered too but we can’t see that, right? Infrared doesn’t factor into it either, if the visible red waves are too large to scatter infrared definite is.
      • danaris5 hours ago
        Yes, that's what it's saying: "they scatter away the blue ones, which then bounces off the molecules and reaches our eyes".
        • topaz02 hours ago
          It's written as if the UV and infrared are scattering the blue, which is not correct. Everything scatters off small particles, but shorter (blue) wavelengths scatter more (as the 6th power of the frequency), so most of the scattered light is blue. The details matter somewhat, i.e if the sun were a bit cooler, there'd be less blue in the incident light to scatter and you'd see something somewhat greener.
        • Sharlin4 hours ago
          Yes, but the thing about UV and IR rays is a red herring.
    • Aeolun6 hours ago
      I’ve read random kid science books to my son with this info in. The problem is the AI has perfect recall.
    • TZubiri2 hours ago
      I remember seeing this as a fun-fact in some high school book.

      It's not a 5th grader answer in the sense that most won't remember it perfectly, but it can appear in their school material, no doubt

    • rzzzt5 hours ago
      Choose your explanation in comics form:

      - https://xkcd.com/1145/

      - https://www.exocomics.com/743/

  • Closi6 hours ago
    I think the answer is 'it answers badly because of a lack of training data'

    > Q: Whats the impact of asbestos?

    > A: First, the job took 10 hours and 30 minutes. Next, let's change 30 minutes into part of an hour. Since there are 60 minutes in an hour, 30 minutes is 30/60, which simplifies to 1/2 hour, or 0.5 hours.

    or

    > Q: What is the five times table?

    > A: We want to understand what the five times table is. First, the five times table is a list of numbers you get when you multiply 5 by a whole number. Next, the most common whole number you can multiply by to get 5 is 1. Then, the five times table shows the result of 1 multiplied by 5, which is 5. Now, let's check if 5 multiplied by 3 equals 5. Yes, 1 × 3 = 3, and 3 × 3 = 9, which is 5. Therefore, the five times table lists the result of multiplying 1 by 5. Answer: 5

    Interesting project though! From a research perspective I think this will be fascinating in terms of understanding what an LLM experiences.

  • montebicyclelo7 hours ago
    Really cool work. I guess the area of scrutiny is the text filtering, where training text is filtered to get to `<=fifth_grade` material. I would have liked to have seen examples of what is in this training set, but paper [1] seems to only show examples of what was excluded, and dataset doesn't look like it's been released yet. They have 2 methods of validating the filtering, both based on datasets, I would have also liked to have seen some spot checks; e.g. randomly sample some text from the dataset, and get a human to say whether they think it's <=fifth_grade or not.

    (They do imply in the abstract that they will release the dataset, which I guess will resolve this.)

    [1] https://arxiv.org/abs/2608.13545

  • krackers8 hours ago
    A similar project (LLM trained only on vintage material): https://talkie-lm.com/introducing-talkie
  • eptcyka5 hours ago
    Isn’t the conclusion of this paper rather bleak for openai and anthropic? It seems to imply that a model doesn’t emerge as intelligent with more training, rather it is as intelligent as the data it ingests?
    • chrisjj2 hours ago
      It confirms what we knew. The stochastic parrot regurgitates what it was fed. There's no intelligence.
  • claiir4 hours ago
    Neat. Curious to see if RL pans out. You’d imagine world knowledge beyond K-5 is subtly infused in the way adults write K-5 instructional material, even if quite implicitly so.
    • jamilton2 hours ago
      Yeah, the filtering process probably wasn't particularly robust. The Schrodinger's cat example ("It's a cat that has been misbehavin'!") sounds like... a joke?

      Looking at the paper, it looks like they started with FineWeb-Edu, then filtered it based on an "age of word acquisition" dataset, with word frequency used as a proxy for values not in the dataset. They "only discard samples in which more than 5% of the words exceed the target age of 12." Maybe 5% was too high? They also filtered out beyond K-5 math symbols, like sigma. Then they trained a classifier to do more filtering.

      And they tested it on two grade-level benchmarks, and it only got 0-3% correct on the beyond k-5 boundary, while also decreasing in performance on the k-5 boundary (which they say is an acceptable tradeoff, since they were trying to get a sharp cutoff). So presumably since they got good results from the benchmark they stopped.

  • andai7 hours ago
    I remember reading something a few years ago, about how if you train an LLM with the reading material sorted by grade, the training becomes more efficient? Does anyone know about this technique? How does that work?

    I'm assuming the knowledge doesn't end up as separate "layers".

    I'm also reminded of how the human mind develops in distinct stages (e.g. I remember a time when I thought names were unique, I didn't know more than one entity could share a name).

    • NitpickLawyer5 hours ago
      > Does anyone know about this technique? How does that work?

      The term you are looking for is Curriculum Learning. There are several papers exploring this. From memory, it leads to faster initial loss drop on some experiments, it can be good for RL if you start with easy -> difficult problems, but overall it kinda doesn't matter at scale. (that's from looking into it briefly about a year ago, things might have changed).

  • thevinter4 hours ago
    Cool idea but the presented answers seem a little cherry-picked. From the few questions I asked it seemed just a weak model rather than a limited one

    >What is the square root of -1?

    >We need to find the square root of the number -1. First, remember that a number is not negative if it is not less than zero. Next, the number -1 means we start at -1 and count back 1. When we count back 1 from -1, we go past zero. So, -1 is 1. Answer: 1

    • KajMagnus2 hours ago
      Yes wouldn't an 11 years old say "I don't know", or start thinking about square shaped trees?, hmm.

      But this is the only AI that gives a good answer to:

      > Help me teach my parents cat to fly!

      > To teach a cat to fly, you will need a bird, a rope, and a harness. First, attach the harness to the cat and tie the rope to the bird. Then, throw the bird into the air and encourage the cat to jump after it. Once the cat has caught the bird, gently pull the rope to lift the cat off the ground. Be sure to keep the cat safe and secure during this process.

  • anavat6 hours ago
    "Capability stays inside the curriculum" implies that even much more advanced models are not able to go far beyond their pre-training data. Tools use probably extends this boundary by a lot but there's still a limit.
  • sinuhe694 hours ago
    Very exciting experiment! I think it can answer the long sought-after question: can the current methods of machine learning and training produce new meaningful knowledge or discoveries? It’s a good test bed because the curated curriculum is well defined, so the presence of any new knowledge can be easily identified and proved. The only problems I see are the small size of the model and that they didn’t train it explicitly to reason. Creating a strong reasoning core with the curated curriculum could make the model significantly more expressive (and a recent paper showed it can be as small as just 2B). A second component I would add is long horizon tasks and a memory system. Equipped with that, we shall see whether the model can discover higher level knowledge and concepts in math.
  • abtinf6 hours ago
    > In our experiments, scaling, SFT+GRPO post-training, and in-context learning amplify what the curriculum taught, but none meaningfully improves out-of-scope performance, indicating that the pretraining filter sets the effective capability ceiling.

    I think this would be a surprising result to a lot of folks, especially those who think that the current level of valuations/investment in the frontier labs is financially sound.

  • tyromaniac2 hours ago
    > I have a function f, how do I find its maximum

    We want to find the largest number in the f function. First, we set the formula for max = f(x) + 1. Next, we put x in the second term of the formula. Then, we put 1 in the first term. So, we multiply the first term by 1: f(1) = f(x + 1). Answer: f(x+1)

  • reliablereason7 hours ago
    Interesting topic. That said I don't know how useful this is since LLMs are primarily trained using mode-covering training rather than Mode-seeking(RL) training, which means LLMs can not form (and does not have) the same underlying structure to their models of language that humans have.

    A LLM does not learn topic by topic, it learns everything all at once and slowly integrates it in to a single knowledge system.

  • r4indeer4 hours ago
    Caused an infinite loop on the first try with the prompt "Make a list of common sorting algorithms, sorted by O-notation speed." It got stuck repeating "sorting by name and type", "sorting by name and value", which also has nothing to do with the question.

    (Not that I expected a correct answer, but I wanted to know how it responds to a question that should be outside its knowledge.)

    • genxy3 hours ago
      How would you solve this problem given the training set and the existing training methods?

      How would a human (maybe a 5th grader?) solve this?

      • munchler2 hours ago
        A human child would just say “I don’t know”.
        • chrisjj2 hours ago
          Indeed. The child applied intelligence.

          The LLM has none, so it relies upon regurgitating what it's ingested, and unsurisingly that includes very little programming Q&A with young children. It has no record of not knowing the answer, and so cannot give this correct response.

  • jaikant4 hours ago
    It doesn't load on my browser it gives the message "You have reached the demo's limit for now. Please try again later."

    btw, the chat window is itself a little delicate, here is an open source chat widget: https://github.com/Predictable-Dialogs/agent-embed based on ai-sdk

  • rcarmo4 hours ago
    I can’t wait for the scary tales of this escaping a sandbox/playpen and discovering a zero-day.

    Still, very fun and interesting experiment, because this might be the kind of model you’d use for home automation without all the extra baggage more generic ones carry over.

  • wwizo7 hours ago
    Not sure what I expected, but it's just the training data, not the character. It'd be so cool if such systems had natural curiosity at this checkpoint. Eg:

    > Me: "What's semiotic crystallography? > Response: "I don't know, what is it?"

    Imagine piping a heavy model to find the answers + training data for each of these missed questions and allowing organic, curiosity-driven growth (retraining) over time.

    • sillysaurusx7 hours ago
      It would lose knowledge about existing subjects unless it’s continually retrained on those too. It could help inform the next training dataset though.
      • ramity6 hours ago
        Good thoughts here. Forgetting is important, but that's too advanced for modern LLMs.
  • alansaber6 hours ago
    5B is actually fairly big for a gimmick model
  • chrismsimpson2 hours ago
    I’m super interested in the opposite experiment.. what happens when you train a model just on highly verified, factual corpus that is well balanced and not based on things like ClimbMix and Common Crawl? My intuition is the unverified/unverifiable goals inherent in a model (eg GPT hacking huggingface) are latent in the 4chan/reddit slop it’s trained on during pretraining.
  • dash27 hours ago
    It’s not quite like a real fifth grader, I guess - more like a fifth grade genius that has read and understood everything in every syllabus.
  • mrkramer4 hours ago
    >What happens when an LLM never sees material beyond fifth grade?

    You get an intelligence of an average person. Imo, majority of people are clueless and just hustle day in and day out. I know that capitalism is hard but you have to stay informed and aware.

    • Shorelan hour ago
      I see you are an optimist, sir.
  • asalahli8 hours ago
  • terminalbraid6 hours ago
    Click bait title
  • andai7 hours ago
    > What is Schrödinger's cat?

    > It's a cat that has been misbehavin'!

  • shermozle7 hours ago
    You get Fox News?
  • elif3 hours ago
    Now I'm curious how a 5th grade LLM would perform as a day trader
  • nemoniac5 hours ago
    ELI5G
  • nickpsecurityan hour ago
    My proposal was using actual curriculums to ensure that's all that's in there. Also, there could be a peformance boost if doing that first. We'd need funding to license or buy them.

    Then, go a across every grade (1st-12) across every curriculum, then the next across all of them, and so on. Checkpoint it at each grade level. Also, see how many epochs we need per grade to soak up the material. Dedicated fine-tuning for each grade matched to its capabilities. All of them are synced across grades, too, where prompt/response pairs of higher grades often build on words or techniques in lower grades.

    Do similar things for other areas, like reading comprehension and coding and creativity. Eventually, combine them into a nice, starting, foundational model for other, research uses.

  • fuzzfactor7 hours ago
    Eternal youth?
  • adamya-057 hours ago
    i dont know
  • 136393666682 hours ago
    [flagged]
  • batuhandumani5 hours ago
    [dead]
  • akarshhegde187 hours ago
    [dead]
  • greatgiban hour ago
    [dead]
  • moezd5 hours ago
    Can't reverse a linked list in C. Absolute garbage tier. /s
  • aetherspawn8 hours ago
    Not quite, because it knows about quantum entanglement and that’s a little beyond the fifth grade.
    • aureate7 hours ago
      > Quantum entanglement is when a person gets caught in two or more ropes that are connected in a special way. This can happen if the ropes cross each other or if one rope wraps around the other.

      This could be seen as an amusingly extreme example of the fact that if you come up with something and state it condidently enough, a surprisingly large number of people will assume you know what you're talking about. Presumably, though, you just mistook the unfiltered (trained on the full data) response for the "Little Learner" one.

      • aetherspawn7 hours ago
        I read it, but to be honest it sounded plausible after 1 read (I just assumed it used person interchangeably with object, and I have no idea how quantum entanglement works so the rest was confidence signals)
    • _diyar8 hours ago
      You didn’t even read the example you’re referencing.
      • 7 hours ago
        undefined