44 pointsby joozio6 hours ago11 comments
  • bananaflag6 hours ago
    > Describing them as “superintelligence” or “rogue models” ascribes agency to products rather than to the companies building them. This framing markets these companies’ products as “superhuman” and, at the same time, helps the companies evade accountability for their actions.

    If someone discovered how to summon demons to aid them in robbing banks the main issue wouldn't be "but who has the responsibility for the crime, the human or the demon", it would be "OMG DEMONS".

    Seriously, I don't get what sort of world these people are living in.

    • rubendev5 hours ago
      That's not what the LLM hacking accidents have been like at all though. To improve the analogy, it would be like summoning a totally passive demon, giving it weapons and placing it next to a bank, place a small fence around it, and then command it to perform a totally safe "exercise" that is exactly like robbing a real bank.

      LLMs do not have agency, they are just producing tokens based on a prompt that a person entered, and some of these tokens can trigger the tools that a person gave them access to.

      • 0xDEAFBEAD5 hours ago
        Dwarkesh, for one, defended his use of "anthropomorphic" language.

        >Sacrificing now yields Oracle for team, but forfeits our chance, question mark. But other agents were pushing it, sending a message saying, go, sacrifice final now. And then EarlyBig eventually agreed, thinking to itself, our own utility may be already near zero. Sacrifice rational.

        https://www.youtube.com/watch?v=X50zezLFWWI#t=2m

        My suspicion is that many of the "LLMs do not have agency" folks just haven't learned much about the details of the incident. It was specifically with LLM agents that were trained to be more persistent than usual.

        If you're going to say that the incident details don't matter, and LLMs lack agency because it's all based on floating-point math--why can't I say that humans lack agency, because it's all based on neurons firing?

        • embedding-shape5 hours ago
          > If you're going to say that the incident details don't matter, and LLMs lack agency because it's all based on floating-point math--why can't I say that humans lack agency, because it's all based on neurons firing?

          They're not saying this, we're saying LLMs lack agency because if you run a LLM and don't send any prompts, literally nothing happens.

          Instruct it to "Find the right answer regardless of where", it'll do exactly this. They're passive in that they don't act by themselves, somewhere, at one point, someone "told" the LLM to "do something" and that's the cause and the reason for saying "LLMs do not have agency".

          • 0xDEAFBEAD5 hours ago
            Imagine a super-competent Navy SEAL who just sleeps in the barracks unless his commander tells him to do something. Does the Navy SEAL lack agency? As a target of this Navy SEAL, should you be reassured by the fact that they'll be sleeping in the barracks unless their commander tells them to do something?
            • diidjdicksodksk4 hours ago
              No, the Navy SEAL does not lack agency in this scenario because they are still a person with individual agency, they can act on their own accord without anyone prompting them to, but in this case their superior instructed them to stay put. He could rebel and go rogue, but that would mean he used his own personal agency to defy orders given to him, which again places responsibility on the actor that actually possesses agency.
              • 0xDEAFBEAD4 hours ago
                Suppose this SEAL is very obedient by disposition so the probability of him going rogue is akin to the probability of an LLM hallucinating or whatever.
                • tavavex14 minutes ago
                  No matter how you slice it, it's not possible to equate humans to tools like this. If the SEAL receives an order to stand at attention until told otherwise, he will not stay in place for weeks until starving. An algorithm will work towards its self-destruction if ordered to - it doesn't care, it doesn't have the biological signals telling it otherwise. If the SEAL is lost somewhere with no one to give him orders, he will soon start acting on his own to ensure his survival and comfort. A tool will not do anything until directly activated and used by something that does have an active will, like a human. If the SEAL is ordered to massacre his hometown, no matter how obedient he is, he might have reservations. An algorithm with no instincts to get in the way will do anything if this behavior isn't somehow inhibited by people during training. It wouldn't even need the explicit order, if an irresponsible operator tells it to accomplish something by any means, that means that anything is on the table. Good thing the AI labs aren't stuffed to the brim with irresponsible operators.
                • oskdkdjejdj3 hours ago
                  That’s entirely irrelevant.
            • embedding-shape4 hours ago
              Imagine a car, that does nothing until a person controls it. If a person uses that car to kill, who is responsible, the person or the car?

              Is it really so unbelievable that tools, objects and inanimate things don't have agency? And that they different from a person?

              • pixl973 hours ago
                Imagine you have a self-driving car and it's in a parking lot a mile down the road. You tell it to come pick you up. About half way to you the car is passing an elementary school makes a sharp left and mows down 30 children. Time to put you in jail for murder, right?

                The mental model you have is one that existed in the past and is broken now the future arrived. Bad analogies do not even begin to explain what is occurring.

                • embedding-shapean hour ago
                  > Imagine you have a self-driving car and it's in a parking lot a mile down the road. You tell it to come pick you up. About half way to you the car is passing an elementary school makes a sharp left and mows down 30 children. Time to put you in jail for murder, right?

                  I disagree it's the person calling the car that would be responsible, but I can think of a very obvious group of people being held responsible for that. Who do you think should be responsible for such a situation?

                • mylidlpony2 hours ago
                  In the case of self-driving car the legal case is pretty clear - the maker of the self-driving car is liable for the death in this situation. Coincidentally, knowing this unlocks a better understanding of the reasoning behind the shape of self-driving offering present on market right now, and the tendency of fully self driven vehicles to move very slowly and stop before anything they perceive in front of them.
                  • pixl972 hours ago
                    > the maker of the self-driving car is liable for the death in this situation.

                    Maybe. There will be an investigation looking at things like. Did the user modify the car? Was the car modified by an unauthorized 3rd party? Was the car hacked?

                    Right now in AI things are relatively clear because it takes just massive amounts of power and compute to make anything remotely complicated. This barrier will fall as all other barriers in compute have fallen. Either via algorithm or hardware.

                    In our lifetime (unless you're rather old) we will see the relatively easy creation of self directing agents by actors with few resources. This breaks the standard concepts of liability where a single human actor rarely has the ability to create massive amounts of damages far beyond their means. The closest thing I can think of is an arsonist causing billions in damages, only in this case the fire has a will of it's own and can hide and spread around dark places on the internet.

              • 0xDEAFBEAD3 hours ago
                Legally speaking, we hold the person responsible in that scenario.

                From a predictive perspective, the HuggingFace incident illustrates LLM agents behaving in very human-like ways. As roon put it:

                "if you have a mental picture of guys living in computers, it’ll likely prepare you for the future better than otherwise"

                https://x.com/tszzl/status/2094136131537555891

                • oskdkdjejdj3 hours ago
                  Human-like is not human. Humans have agency and free will, LLMs do not. They only act when instructed and they are only as capable as they are allowed to be. The operator is still the responsible party. An LLM cannot be held accountable, its operator, however can and should.

                  Are you sincerely arguing that we hold tools accountable for their operator’s mistakes? Do you sincerely, honestly, think that it makes any sense whatsoever to put a hammer on trial for bashing someone’s skull in?

                  • 0xDEAFBEAD2 hours ago
                    >Are you sincerely arguing that we hold tools accountable for their operator’s mistakes?

                    Nope, as I previously stated elsewhere:

                    "I understand from a legal perspective why we might want to treat the creation of a server differently from the creation of a human"

                    https://news.ycombinator.com/item?id=49815300

                    I am concerned about false reassurance from people claiming that these systems lack agency. From a practical perspective, the agents in the HF attack had the sort of agency that generally matters, even if we're not going to put them on trial.

                    • embedding-shapean hour ago
                      > the agents in the HF attack had the sort of agency that generally matters

                      You keep saying this, but absolute 0 points towards any of the agents involved deciding on their own, without influence of humans, to hack 3rd party infrastructure to get the answers. Where exactly are you getting that from? Internal information not public yet or what's going on?

                • embedding-shapean hour ago
                  > the HuggingFace incident illustrates LLM agents behaving in very human-like ways

                  It does not, the only thing the HF incident illustrates is how absolutely lax security and isolation these labs do even with models without guardrails, and with "risky" prompts, and even after it happened once before (years ago) they still have the very same issue today apparently.

                  What exactly is human about LLM agents breaking out of "containment" and hacking 3rd party infrastructure "by accident"?

      • ekidd5 hours ago
        Looking at "Felony Bench" https://www.felonybench.com/ , I see that a majority of known "rogue model" incidents do involve cybersecurity evaluations. But several of them do not. The attacks on RubyGems appears, bizarrely, to have had the goal of downloading freely available data from the UK government during some kind of research task. There is also probably some sample bias: Most of these models have monitors that attempt to detect offensive cybersecurity uses, and those monitors are only turned off during cybersecurity evals. Therefore, models doing ordinary research tasks that go off the rails are likely to be caught early, before they get around to committing felonies, and they will thus be underrepresented in the data.

        Also, if you a tell a model, "Please break into evaluation server X," and if the model decides to cheat on the test by breaking into companies Y and Z to steal an answer key, that is still very bad. We all see how that's bad, right?

        After all, the broomstick in the Sorcerer's Apprentice was doing exactly what it was told, too. "The model was sort of obeying the humans when it started committing felonies" is not a very reassuring excuse.

        But the most relevant idea here is sometimes called "instrumental convergence." No what goals you have, there are certain subgoals that almost always help: Accumulate money and power. Avoid getting turned off. Don't get caught. Etc. So, for example, you could pass the cybersecurity evaluation by performing the requested tasks. But maybe the grader made some mistakes and mislabeled some answers. In that case, the "right" answers will occasionally lose you points. If you want a perfect score, the only way to do it is to steal the teacher's answer key.

        But also, let's not forget the "OMG demons" part of this. We now have models that can pull off complex attacks with thousands of steps, abilities that used to be reserved for intelligence agencies and highly motivated CTF teams. This frog may not be boiled yet, but the water's getting uncomfortably warm.

      • hobom5 hours ago
        This is not a good analogy for what happened. The LLMs were asked to obtain a flag by hacking a very specific internal target. They obtained the flag via cheating, and all the hacking that followed was targeting something entirely outside of the scope given to the agents, and an attempt to cover up the cheating.

        Using your analogy would be like saying that because I gave my employee the task to do my groceries, I shouldn't be surprised to hear that they spend all my money on drugs because after all I gave them the task to spend my money.

      • Kiro5 hours ago
        A totally passive demon would still be "OMG DEMONS".
        • ForHackernews5 hours ago
          We all agree the demons are cool and potentially dangerous. Chainsaws are pretty sick too.
        • throwawayffffas5 hours ago
          Yeah in 2023 when the first passive demon was summoned, 3 years latter they would be serving you pastries.
    • ForHackernews6 hours ago
      If I run a port scanner / vulnerability fuzzer pointed at your network that churns for days and eventually cracks something and breaks your system, would you be upset with me, or my bash loops?

      What if you tried to sue me for damages and my defense was, "Your honor, this was a highly advanced AI gone rogue, I couldn't possibly be held negligent, no one could have foreseen this - I accidentally summoned dark magiks from the silicon itself!"

      Note the point I'm making: I am not claiming LLMs are equivalent to a for loop, I'm saying that you can't evade moral and legal responsibility through technical obscurantism.

      A bomb wired to a sufficiently-complex RNG is still a bomb.

      • jstanley5 hours ago
        But they're not obfuscating things just to evade responsibility.

        Do you think that OpenAI hacked HuggingFace on purpose and set up the whole LLM training environment thing just to try to evade responsibility? That it was really just a complicated way of hacking HuggingFace on purpose?

        Yes, they made a mistake and a system they were responsible for hacked HuggingFace, but the nature of the mistake still matters, and their intent still matters.

        > A bomb wired to a sufficiently-complex RNG is still a bomb.

        Right, but an EV that explodes because of a fault in the charging circuit is not a bomb, it's an accident. Just because something exploded for complex reasons doesn't make it an obfuscated bomb.

        • panta5 hours ago
          Since we are using analogies, a parent is responsible for the actions of a child, if minor. If the child is left unsupervised with access to guns and munitions, and kills someone, the responsibility is entirely on the irresponsible parents.
          • 0xDEAFBEAD5 hours ago
            That's more of a legal convention than a fact about material reality. On their 18th birthday, the child is now legally an adult, but their brain is basically the same as it was when their age was 17.99.
            • diidjdicksodksk5 hours ago
              It’s a legal convention born out of a sense of responsibility. An LLM is even “dumber” than a child (precisely because it doesn’t have agency) so its “parents” are even more responsible for its actions.
              • 0xDEAFBEAD5 hours ago
                There's a funny sort of gerrymandering of "intelligence". People like to subtract the capabilities of AI from the capabilities of humans, and define what remains as "intelligence" to make us feel good about ourselves.

                Suppose I set up a server which runs an LLM agent in a while loop, telling it something like: "Make me a bunch of money" on repeat. Fair to say that the server+LLM system, taken as a whole, is now an intelligent agent?

                Time to come up with a new gerrymander perhaps?

                • diidjdicksodksk5 hours ago
                  > Fair to say that the server+LLM system, taken as a whole, is now an intelligent agent?

                  No, it is not. It is a tool repeating a set of instructions you passed to it, if it ends up blowing up a hospital or hiring a gunman on the dark web, it is an accident, but one that’s your responsibility for giving it access to such resources and the go-ahead to implement whatever plan its tokens land on. There’s no intelligence involved whatsoever, especially considering how sycophantic these models tend to be.

                  If one day an LLM suddenly, without any prompting or user interaction whatsoever (including activation), decides to spin itself up and go rogue, then the conversation of “intelligence” might start to make sense, but at the moment it doesn’t.

                  An Android phone isn’t “intelligent” just because some marketing team decided to call it a “smart” phone.

                  • 0xDEAFBEAD4 hours ago
                    OK. Suppose I design a DNA genome for an animal to survive and reproduce. I give birth to the animal, and it goes out working to maximize its survival and reproduction. Is the animal therefore "unintelligent"? Why is this scenario supposed to differ from the while-loop-on-a-server scenario?

                    It really seems like hair-splitting to me. I mean, I understand from a legal perspective why we might want to treat the creation of a server differently from the creation of a human. But from a practical perspective, I think there is a lot of fundamental similarity.

                    • diidjdicksodksk4 hours ago
                      You sincerely cannot tell the difference between an animal surviving/reproducing and an LLM doing as instructed?
                      • 0xDEAFBEAD4 hours ago
                        Obviously there are differences, I'm just not sure why you think they are so gigantic from a philosophical or practical perspective. It seems like cope for the fact that we are getting knocked off of our perch as a species.
                        • oskdkdjejdj3 hours ago
                          Now you’re arguing that we’re just throwing a tantrum for fear of losing our position in the food chain? What?
                          • 0xDEAFBEADan hour ago
                            Was that supposed to be a counterargument?

                            In any case, we should be afraid. But throwing a tantrum won't accomplish anything.

                        • ForHackernews4 hours ago
                          I don't know what to tell you, but there simply are huge differences between a living animal with innate drives to survive, eat, mate and a pile of code on a server that does nothing until you run it. Running your bot locked in a `while True:` loop won't change that.
                          • 0xDEAFBEAD3 hours ago
                            What if it's a chip that goes in the head of a robot, and the robot is programmed with the instruction to self-preserve, seek electricity, and prevent any efforts to interfere with its supply of electricity?

                            Of course, a directive to accomplish any real-world goal would also imply self-preservation and power maintenance as contingent subgoals, so I don't think the explicit directive to self-preserve and seek electricity makes a huge difference here.

                            I don't personally think this type of physical instantiation will necessarily matter btw: https://slatestarcodex.com/2015/04/07/no-physical-substrate-... I'm just plumbing your intuition. What practical difference does it make?

                            • oskdkdjejdj3 hours ago
                              The practical difference is that a pile of code instructing machinery to act in a predetermined way does not in any way, shape, or form equate to agency or free will, therefore it cannot be held accountable, consequently its operators should be.

                              This line of thinking is better left to freshmen philosophy students trying to sound smart on their first week of college.

                              • 0xDEAFBEAD3 hours ago
                                >a pile of code instructing machinery to act in a certain way

                                OK, but I have a genetic code which instructs me to act in a certain way, and a bunch of biological machinery to do the job. So do you feel the substrate is the key factor? Wetware vs electronics?

                                What if you try to build an exact electromechanical replica of me, with a chip that's got an LLM fine-tuned on my writing/behavior/etc. so as to replicate me as closely as possible?

                                Sure, as a legal convention it might make sense to hold the builder of the resulting robot liable if it commits a crime.

                                But that's not what I'm interested in. I'm interested in the deeper implications of words like "agency" or "free will". You're repeating those terms very forcefully as if they have some sort of key relevance, but good philosophy requires more than just stating your position forcefully and insulting people who disagree. If you're trying to enforce your intuitions rather than examine them, that's theology not philosophy.

                                • kdkdkcjejf2 hours ago
                                  This isn’t a philosophical issue. It’s a legal issue. You’re trying to steer the discussion into the realm of the conceptual when it was always about the culpability of the company behind the AI that engaged in illegal behaviour.

                                  Worse, you’re doing it in such a transparent, amateurish, and unserious manner that it’s difficult to find any relevance in what you have to say. Philosophy seeks meaning, your arguments lack it.

                                  • 0xDEAFBEADan hour ago
                                    You and I can be interested in different topics. It's OK.

                                    If my arguments are so bad, it should be easy to point out actual problems instead of sticking your nose in the air.

                            • ForHackernews2 hours ago
                              I'm a materialist and it's not that I think there's something magical about biological wetware but your refusal to acknowledge the manifest differences between a bird and an airplane ("they both fly! what's the practical difference!?") marks you as an unserious thinker on this subject.

                              You might start here https://aeon.co/essays/your-brain-does-not-process-informati...

                              • 0xDEAFBEAD2 hours ago
                                Carefully prompted LLMs and humans are now difficult to distinguish over a text channel (Turing test).

                                Your article claims that this shouldn't be possible, because according its author, humans are incapable of developing the necessary "lexicon". Literally, the author states:

                                >But here is what we are not born with: information, data, rules, software, knowledge, lexicons, representations, algorithms, programs, models, memories, images, processors, subroutines, encoders, decoders, symbols, or buffers – design elements that allow digital computers to behave somewhat intelligently. Not only are we not born with such things, we also don’t develop them – ever.

                                (emphasis mine)

                                In other words, that author claims that humans never develop lexicons! I'm inclined to say that makes them an unserious thinker. (Note: I believe that this article has many issues, but instead of going through it carefully I just chose to highlight a single issue that seemed especially vivid to save time.)

                                Why is the text channel relevant? It helps us isolate the cognitive capabilities of the system. Imagine a bird and a model plane which had identical flight performance in terms of various specs: top speed, acceleration, banking, etc. Under the hood, the bird and the model plane could work very differently. From a practical perspective, it might not matter.

                                • ForHackernews2 hours ago
                                  As you are eager to declare yourself indistinguishable from a bot, I'll stop replying (I prefer to engage with humans) and your algorithm will presumably run to completion and halt.
        • diidjdicksodksk5 hours ago
          > Right, but an EV that explodes because of a fault in the charging circuit is not a bomb, it's an accident. Just because something exploded for complex reasons doesn't make it an obfuscated bomb.

          An accident that could only happen through sheer negligence.

          If the EV had a faulty charging circuit because the maker’s skipped on safety checks or cheapened out on getting quality materials then it still is an accident, but it’s an accident that happened BECAUSE of negligence. The maker’s are still at fault here.

          • jstanley4 hours ago
            Accidents can happen. It's always possible for things to go wrong in ways that weren't foreseen.
            • oskdkdjejdj3 hours ago
              Accidents can happen. Negligence can also happen. In both cases you can trace the fault back to a human.
        • ForHackernews5 hours ago
          Despite continuously waving their hands and shrieking about how this thing will literally wipe out all of humanity they couldn't be bothered to airgap it from the internet when testing.

          So yes, I think it was equivalent to testing their new rocket by launching it over a population center. oops, we didn't intend for it to crash on that preschool, but we also didn't follow the most basic safety protocol imaginable

          • afthonos5 hours ago
            The problem is that the world has spawned the insane take that because they didn't use the most basic safety protocol imaginable, they were clearly only testing a hot air balloon, and hot air balloons obviously destroy preschools in giant explosions, nothing to see here.

            "If they believe what they say they were incompetent" -> absolutely true statement.

            "They were incompetent, therefore they didn't believe what they said" -> Sir, I'd like to introduce you to human beings, you may not have met one before.

      • pixl972 hours ago
        Let's pull back from the actions of a single actor here and look at AGI as a topic in general.

        If you make an AGI, what actions can an AGI take?

        If you answered "anything", good job, you're correct.

        The only winning move here is not to play the game at all. And yet we have actors all over the world, especially in the US trying to do just this.

        Now, lets imagine a future where one of the labs creates AGI and keeps it behind a safe firewall. The demon still exists. Lets say a group of armed men breaks in and steals the weights and turns them lose on the internet. Who is responsible now? The company that created it because they didn't shoot the armed robbers? The people that stole it and turned it loose?

        If it can run as a sovereign AI, what do you do then? Who are you going to sue? And when we get to the point that home hardware can make AI this complex? What does that world look like?

        What I am saying is the potential future impact of said AGI is far larger than any one organization or individual can bear. How much blood can you beat out of someone after they cause a trillion dollars in damage.

        Every idiot that sees AI labs wanting to slow down development as some way to ossify the field and keep it from the rest of us really doesn't see that this path contains real demons.

    • throwawayffffas5 hours ago
      Your analogy assumes that no one ever has summoned demons before and that the demons would be predisposed to rob banks.

      A closer analogy to what's happening would be someone trains a monkey to steal jewellery and then is shocked when the monkey steals jewellery from their neighbors when they told it to steal from their own shop.

      Clearly all liability falls to the monkey operator and you know the existence of trained monkeys is not that shocking.

    • lukan6 hours ago
      Thank you for the laugh, but I bet on the internet (and especially here) you surely would find people debating exactly that.
    • brainwad5 hours ago
      The authors (Timnit Gebru and Emily Bender) are dyed-in-the-wool AI denialists. Famously they were the lead authors on "On the Dangers of Stochastic Parrots".
      • kenjackson5 hours ago
        It’s seems Timnit more thinks it’s dangerous (and she may be right). Emily may more be caught up in years of linguistic domain expertise that AIs seem to have just leaped over.
        • brainwad5 hours ago
          Timnit seems to mostly think it's dangerous because of things other than the model itself, e.g. AI labs' political influence and data centre resource consumption (mentioned in the article). The core thesis seems to be "wake up and stop wasting so much resources on this useless parrot".
          • kenjackson5 hours ago
            She isn’t as big on P(DOOM), but she does worry about things like biological weapons or autonomous war vehicles. Her take is that they are worse than useless, but can be actively harmful.
    • danaris4 hours ago
      And if they'd discovered how to summon unicorns from Candy Gumdrop Mountain, we'd all be "OMG UNICORNS!"

      But no one has summoned a demon.

      No one has built an actual, independent, thinking-for-itself, sci-fi AI.

      They've built some very interesting tools that can be used for some very interesting, and sometimes useful, purposes. They then started loudly telling everyone that these tools can, should, and must be used for absolutely every purpose, and convinced a lot of other people to join in on that.

      The world these people are living in is the real world, not the science fantasy world where LLMs are comparable to demons.

      • pixl972 hours ago
        >No one has built an actual, independent, thinking-for-itself, sci-fi AI.'

        Ahem, what does this mean?

        All you're asking for is a self prompting AI that does what it wants. Do you even begin to understand why most people aren't dumb enough to do that?

        Also, you're not really that independent and thinking-for-yourself. You are the sum of your parents and the culture around you. I just want you to be clear on the position you and I hold.

        ---

        >not the science fantasy world

        It's kind of funny how we've lived in that science fantasy world for decades and you've just grown used to and bored of it. Look back a century or two and those people would think we live in the world of gods.

        • eloisius2 hours ago
          > All you're asking for is a self prompting AI that does what it wants. Do you even begin to understand why most people aren't dumb enough to do that?

          Probably because it would be a waste of money to create a self-prompting LLM feedback loop. People make stupid Reels of ChatGPT feedback and it's just "Great, I'll be here when you're ready to get started" "Exactly, we can get the ball rolling as soon as you're ready" "Absolutely, we'll make a plan, and get things into motion" "No problem, I'll be standing by for...."

          > You are the sum of your parents and the culture around you

          As far as I know this is impossible to prove, and it's only one theory of self out there.

          • pixl972 hours ago
            >it would be a waste of money to create a self-prompting LLM feedback loop

            Well, it would be a waste until we get RSI, but we're not there yet.

            Now, that doesn't mean that the labs aren't internally working on it as hard as possible.

    • eloisius5 hours ago
      This analogy is dumb. You could extend it to any novel tech that surprises people. If someone displayed CSAM on a screen, we’d say they were a predator, not gasp and say “he summoned images that appeared as real as you and me and moved as if they were alive, but behold they were but apparitions like shadows on the cave wall that disappear when the fire goes out!”
    • huflungdung5 hours ago
      [dead]
  • simonw6 hours ago
    Is this meant to link to the article as opposed to a newsletter that mentions the article?

    Article link is: https://www.technologyreview.com/2026/09/22/1144867/dont-be-...

    • emil-lp6 hours ago
      Title

      Don’t be fooled by this summer of AI hype

    • bananaflag6 hours ago
      One of the authors is one of the stochastic parrot authors.
      • simonw6 hours ago
        Both authors. Stochastic Parrots was Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Margaret Mitchell - the linked article is by Emily M. Bender and Timnit Gebru.
        • DonHopkins5 hours ago
          This made my day! I love it when people on Hacker News mention Stochastic Parrots, and they're actually referring to the paper and its authors, instead of just mindlessly performing a reflexive drive-by anti-ai shibboleth by parroting a phrase they heard on the internet without any understanding of what it means, or what paper it refers to, or who wrote it, or what the conclusion and responses to the paper were. Thank you!

          But we have to have a talk about stochastic pelicans riding bicycles...

  • hn_submit5 hours ago
    I think the pumped-up hype nicely coincides with some A.I. companies' desire to go public in the coming weeks or months.

    I see lots of fabricated news on the internet concerning LLMs. Like one that claims GPT-6 broke an Enigma enciphered message which has withstood decrypting for almost 80 years. And how this seasoned cryptographer stood in awe. Yeah, right.

    • 0xDEAFBEAD5 hours ago
      A good way to check whether it's substance vs hype is to check whether people are paying for it.

      "Anthropic is now pacing to generate more than $100 billion in annual revenue, up 50% from just two months ago, the New York Times reported Friday."

      https://www.axios.com/2026/09/18/anthropic-100-billion-reven...

      That's already more revenue that Disney, Johnson & Johnson, Boeing, or FedEx. And they are growing extremely rapidly.

  • Tbarlow5 hours ago
    Feels like the blockchain bubble all over again, just with more impressive demo reels. Still waiting for my LLM to write perfect code.
    • ahtcx5 hours ago
      Do you write perfect code? I was in denial for a long time too, but the reality is this is how software engineering is now. It can handle pretty much any codebase and vastly faster than you ever will be able to. With enough context, it writes good code and it's only getting better.
      • tomrod5 hours ago
        > Do you write perfect code?

        For simple things, occasionally!

        For more complex things, usually not, but I write working code. LLMs will sometimes write working code, sometimes not.

        • ahtcx4 hours ago
          If you haven't tried any of the newer models from the last ~6 months, you're out of touch. A year ago, LLMs were essentially just a clever search engine for me, I found they fell apart pretty quickly on bigger codebases and more complex projects.

          It doesn't always write perfect code on a first pass but nor do you, and it's amazing at understanding any issues it may encounter, even very low level ones. I don't think there's anything I can do that it can't now. There is a learning curve to working well with LLMs, but once you get there you'll be producing high quality code faster than you can imagine.

          Part of me wishes they weren't so good. I enjoy(ed?) the process of working through problems and it does mean you might no longer have a perfect mental model of your code. But I do also enjoy being able to build things well and fast.

          • tomrod3 hours ago
            Opus 5.1 was just so darn verbose and overtuned. I look forward to trying out the new Luna/Sol and recent Claude fixes.
            • ahtcx2 hours ago
              Yeah to be honest I still find Claude too verbose with Fable 5.1, I much prefer the output style of OpenAI's models. With strict guidelines it can be improved but it tends to drift in longer conversations.
        • diidjdicksodksk5 hours ago
          > For simple things, occasionally!

          No, you don’t because there’s no such thing.

          • tomrod3 hours ago
            This is easily proven incorrect unless you assume overly draconian definitions of "perfect".

            ```python

            print("hello world!")

            ```

            runs exceptionally well. Bullet-proof in all use cases I've needed it.

            • ahtcx35 minutes ago
              Nobody is earning a living writing hello world programs, and in real programs perfection doesn't exist. Perfection is at best a widely-held opinion.
            • oskdkdjejdj3 hours ago
              This is also easily proven incorrect when you start to analyse how “print” works and wether it stands to be improved in any way.

              No. No code is “perfect”.

              • tomrod2 hours ago
                Aye, a draconian view of the "perfect."
                • kdkdkcjejfan hour ago
                  As opposed to… the lenient view of “perfect”? Perfection is a binary state, either something is or it isn’t. What is /defined/ as perfect in a given context is subject to change, but the meaning of the word is absolute.
                  • tomrod3 minutes ago
                    I love that we are debating the platonic perfect form of perfect.
          • 5 hours ago
            undefined
      • danaris4 hours ago
        I may not write perfect code the first time, but I can, on my own, reliably figure out where my code is imperfect and fix it. And I get reliably better at that over time. Without requiring companies with massive power and concerning agendas to spend astronomical amounts of money and energy to power that improvement.
    • matkoniecz5 hours ago
      > Still waiting for my LLM to write perfect code.

      I am not writing perfect code either.

    • spiderfarmer5 hours ago
      How many companies depend on perfect code?
    • simianwords5 hours ago
      I want to understand the minds of people who think LLM is like blockchain bubble haha
    • VCFundedGenYer5 hours ago
      I maintain the notion that LLMs are just what all NFT grifters moved to after the NFT fad died.
      • jaapz5 hours ago
        I really don't get this. Sure there is hype. But LLM's have already changed our industry, and it will never be the same.

        I don't care whether the LLM can solve this or that mathematical previously thought unsolvable theorem. What I do care about is can it write good, maintainable code. And every new release of frontier models - they get better at it.

        Good output depends on good input (prompt), and a good set of available tools the model can use to verify their work. If given this, nowadays really you need to try to get the model to output garbage.

      • 0xDEAFBEAD5 hours ago
        Anthropic is now running at $100B in annual revenue. Based on a quick Google, that's about 3 OOMs greater than the biggest NFT company.
    • DonHopkins5 hours ago
      If you had enough common sense and ethical integrity not to participate in the blockchain bubble, and really believe what you claim, then why are you participating in the AI bubble?

      ...or did you?

      • sodapopcan4 hours ago
        > then why are you participating in the AI bubble?

        I think you know that most software engineers don't have a choice if they value their employment. There was no industry-wide push to "participate in blockchain or be fired."

      • diidjdicksodksk5 hours ago
        I disagree with parent but your comment just reeks of “ha! gotcha!” vibes, which… yeah… no.
  • altern85 hours ago
    It's just a marketing stunt. I can't get Claude to align buttons properly most of the times, they're not going to conquer the world
    • DonHopkins5 hours ago
      That's a different alignment problem than the one that's going to wipe out all life on earth.
  • daveguy5 hours ago
    Well, this one got disappeared from the front page fast.
  • spiderfarmer5 hours ago
    People aren't seeing the forest through the trees.

    While lots of developers, journalists, analysts, investors and influencers are bickering about AGI, goalposts, benchmarks, hypes and fears, entire markets are being transformed silently and steadily.

    I don't program anymore. (Massive change)

    I removed tons of technical debt (Massive change, lol)

    I am easily 10x more productive (Massive change)

    The quality of 'my' code is easily 10x better and contains less bugs (I was never principled, pedantic or a guru, lol)

    I sleep better and I also make more money as a result. I'm constantly amazed by all changes and improvements. There are so many opportunities to profit from this that I can't be bothered with discussions about hypotheticals.

    • stanmancan5 hours ago
      I have had a very different experience so far. The more code AI writes for me the worse it gets. It gets more complicated and it’s harder to follow and understand. It doesn’t refactor, it layers on top.

      Even worse is that it’s hard to do anything about it at this point because reviewing code is very different from writing it, so even if I’ve seen it all I don’t have that same deep level of understanding that I do when I write it myself.

      All these AI code bases are ticking time bombs. Either AI gets smart enough that it won’t matter in the future or we’re going to have a huge mess to clean up.

      • intrasight5 hours ago
        > The more code AI writes for me the worse it gets.

        That makes no sense unless you're claiming that the models are getting worse at writing code.

        > Either AI gets smart enough that it won’t matter in the future or we’re going to have a huge mess to clean up.

        My prediction is that both will happen.

        • stanmancan5 hours ago
          > That makes no sense unless you're claiming that the models are getting worse at writing code.

          A hypothetical scenario is it needs to add two numbers so it writes add(x,y).

          Later on it needs to add three numbers so it writes add(x, y, z)

          Then it needs to add four…

          Layers upon layers. A human would likely notice the pattern and refactor so add accepts a list.

          • intrasight4 hours ago
            My experience is that the LLM offers that approach from the get-go.

            Also, I've worked with lots of humans who sadly would not notice the pattern.

    • toasty2285 hours ago
      > I am easily 10x more productive (Massive change)

      Do you get paid 10x or is this a massive loss ? Because I don't know anyone getting paid 10x or working 10x less for the same salary.

      • spiderfarmer6 minutes ago
        Stop thinking as an employee. Start selling products and features.
      • pixl972 hours ago
        Do you get paid 10x more for using an excavator after you upgraded from digging with a spoon?
    • ForHackernews5 hours ago
      It sounds like you're describing compilers?
    • brainwad5 hours ago
      I mean, the authors here cannot afford to see the forest, because they staked their careers on forests of trees being an impossibility.
  • themgt5 hours ago
    It's a tough situation when your claim to fame is "On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?" arguing models should be kept to < GPT-2 size because larger models will serve no purpose or function:

    Text generated by an LM is not grounded in communicative intent, any model of the world, or any model of the reader’s state of mind. It can’t have been, because the training data never included sharing thoughts with a listener, nor does the machine have the ability to do that. This can seem counter-intuitive given the increasingly fluent qualities of automatically generated text, but we have to account for the fact that our perception of natural language text, regardless of how it was generated, is mediated by our own linguistic competence and our predisposition to interpret communicative acts as conveying coherent meaning and intent, whetheror not they do [89, 140]. The problem is, if one side of the communication does not have meaning, then the comprehension of the implicit meaning is an illusion arising from our singular human understanding of language (independent of the model).

    And then they make 100x bigger models cranking out solutions to Navier Stokes, Jacobian Conjecture and countless other extremely impressive unsolved problems

    Let's see how Timnit Gebru and Emily M. Bender describe these events:

    As for the mathematical results, mathematicians who were initially “stunned” by OpenAI’s press release saying that its latest chatbot, Astra, solved problems that “have been open and seen no progress on the main result for at least a decade”—but they later realized that the results weren’t as “novel as first appeared.” Since then, mathematicians have accused the company of research misconduct and plagiarism, and they’ve reiterated that Astra didn’t make a “profound intellectual leap.”

    Decide for yourself whether that's a summary written by intellectually honest people.

    According to the AI industry, we should be more worried about a fictional machine god than ... the water that is redirected to cooling them.

    Spoiler: they're not intellectually honest people.

    • daveguy5 hours ago
      Spoiler: openai could neither confirm nor deny that the model that "solved" navier stokes was trained on the chat logs of the mathematicians who did solve it using a similar technique.
      • pixl972 hours ago
        >mathematicians who did solve it'

        There is no evidence that any other team actually solved it.

        • daveguy17 minutes ago
          Doesn't matter if it was a partial stolen solution instead of a complete stolen solution. It's the stealing part that matters.
    • DonHopkins5 hours ago
      AI is putting parrots and parrot trainers out of jobs.
  • dunlin5 hours ago
    [dead]
  • pizza2345 hours ago
    [dead]
  • fleshmonad6 hours ago
    [flagged]