Your car is still the same. Your dishwasher is still the same. The train you take to work is still the same and never comes on time. Fuel is more expensive. The roads are still congested with traffic. Your kids are (probably) doing worse at school. Food costs more. Houses cost more. Rent is higher. Buying a computer or a PS5 is more expensive. Politics is still full of mentally unstable people. The environment is getting worse. You will still die of heart disease or cancer. Food quality is worse. Your kids can't find work. Wealth inequality is accelerating.
But hey...on the flip side...a lot of people are rapidly building software (that nobody is using), and AI is solving mathematics problems (that nobody - including mathematicians - wanted it to solve)
Imagine what a "country of geniuses in a datacenter" will be able to do? Apparently...nothing.
So basically you get personal digital assistant - very useful - but not changing my life.
What would I need for real transformation though? Food, house cleaning, commute. Even if we had ASI tomorrow, what could it do to change my life, besides taking away my job? Design good house robot to help me with food preparation and chores... Maybe finally self driving cars launched in my area? Or it will take control, "optimize humanity survival chance" and realize we don't need as many people on this planet.
I'm noticing this even in industry. Yea, we can build and iterate on ideas 100x faster but has this turned into anything meaningful? Not really... Hell, look at the big companies spending billions on inference... nothing shipped and things are still buggy as ever. Where is everything?
> [Your train] never comes on time. Fuel is more expensive. The roads are still congested with traffic. Your kids are (probably) doing worse at school. Food costs more. Houses cost more. Rent is higher. Buying a computer or a PS5 is more expensive. Politics is still full of mentally unstable people. The environment is getting worse. You will still die of heart disease or cancer. Food quality is worse. Your kids can't find work. Wealth inequality is accelerating.
None of this has to do with AI.
> But hey...on the flip side...a lot of people are rapidly building software (that nobody is using), and AI is solving mathematics problems (that nobody - including mathematicians - wanted it to solve)
Plain wrong on both counts.
> Imagine what a "country of geniuses in a datacenter" will be able to do Apparently...nothing.
Nonsequitor, probably based on a personal beef. Datacenters, at best, are understood to be for hosting more AI - not for people to somehow live in as an arcology of discovery.
It reminds me of the Ezra Klein quote that “OpenAl would need permits to cover its parking lot in solar panels, but it can accelerate into recursive self-improvement, as best I can tell, whenever it so chooses."
Videos are shittier. Most songs are even shittier than the autotuned pieces of shit already were. Ads are worse (even though nobody thought that was even possible). Customer support is worse.
> But hey...on the flip side...a lot of people are rapidly building software (that nobody is using)
That nobody is using and that are worse than what we used to have. And because nobody's using them, the countless issues plaguing them aren't even reported by users anymore (there are no users anymore). Full of bugs, usability issues and gigantic security holes (but which doesn't really matter because nobody's using them).
It's not even clear that it's making long-standing software (like say Linux or Emacs or Blender or QEMU) better at all.
Now I'd say though: most of the issues you mentioned have nothing to do with LLMs / AI. But it's very clear that AI ain't solving much of the world's problems and that the promises we got years ago didn't materialize.
This reply summarizes the problem from a different angle.
Things worked better before we started putting bad software in everything. As a software engineer, I know what good software looks like. The fact that almost all software is now bad -- including the software that drives AI and the software that AI generates -- makes me want to flip the table. We know how to write better software and the general public should hang us if we continue to refuse to deliver it.
Has any technology made an impact on physical world faster?
Imagine being in charge of a physical factory, sure you might have a 200 USD Claude subscription but you also have a floor full of physical machines that have, if you're lucky, an undocumented debug interface and a paper manual. Not to mention safety inspections, insurance, etc.
You can't vibe code a factory floor. Not yet, perhaps in 5 - 10 years but there are some very hard, non technical, problems that need to be fixed before we'll see meaningful improvement there.
This is far from being proven as a net gain and there are massive lawsuits incoming related to AI recommending suicide to kids and others.
Safety systems are not optimal but improving everyday and although I doubt we'll ever get a perfect safety system, it seems that we're not very far from a "good" safety system. The days of Google recommending glue on pizza and generating diverse Nazis seem behind us.
Not to mention all the untold good that LLMs have done. Personally I've seen that LLMs are incredibly good at pointing conspiracy nuts and extremists into the right (moderate) direction.
In a polarised world, we definitely need that.
It needs to a potential field with practical, economical, "real life" attraction well. I.e, robotics, real economic efficiency gains, manufacturing novelties.
The worry is that the 1% is attracted towards a non-practical money hole. I.e: Token burn for the lols, sophisticated software systems that dont provide actual value outside of giving NVIDIA cash.
Meanwhile, the average user was still using ChatGPT, which probably felt like it had plateaued over the last few releases. The real power of these models became apparent when paired with a terminal. Agents like Muse and OpenClaw are now attempting to bring that same Claude Code-like power to everyday users through a simple chat UI.
There’s no reason to think this.
There are lots of good reasons to think this.
The first is that LLMs used to be bad at each of these nerd things, then toppled them like dominos. There's a pattern over time.
Another reason to think this is g. Across every known measurement, human (and animal) intelligence is convergent. Being better at one thing correlates with being better at another thing. Although the reason is not perfectly clear, the pattern is well established, and seems to apply to LLMs too - GPT-6 is smarter than GPT-3 at everything, not just math. There's no reason to expect different for GPT-9.
If anything, in some domains frontier models have become worse - claudisms and chatgpt idioms are making them notoriously bad at prose without a considerable amount of prompting and tweaking.
For example, they have become more and more unintelligible when you ask for explanations or descriptive text. They assume you see the same context as them and shortcut explanations.
I've had to craft a skill to get them to produce remotely understandable explanations of even mildly complex/non-mainstream subjects.
This may be Curse of Knowledge https://en.wikipedia.org/wiki/Curse_of_knowledge on their part, and it also impacts human experts but still. Becoming better at one thing does not mean you become better at everything else, although I do agree with you that the better they become the more things there are they become good at, but their ability is still quite jagged and maybe increasingly so.
LOL. Human (and animal!) intelligence is notoriously jagged. We're all basically idiots except for very narrow areas where we focus.
And then using emgui they could expose full fidelity IDE+Agent workflow interface (like DeepSeek Harness and similar) that's also exposed as MCP to be operated by another agent and JSON RPC for over the network but no. Skill issue?
PS: Even full fidelity Photoshop is possible with such high fidelity UI that has CLI, JSON API and MCP server all in the same single binary < 90 MB. For reference, see PhotoCraft or VectorCraft or WordCraft or PdfCraft.
Really? I start with a conversation for maybe 5 turns or so, where I ask it what would need to change, what the API might be, any database schema changes, URL schemes, and so on, and finally ask it to break it down into commits. Then I let it go, implementing 3-10 commits at a time via subagents. It usually gets the UI somewhat wrong, so there are followups to fix it. This is with Sol and Luna subagents.
Is that what other people see?
We have already reached "peak LLM" in terms of what normal people realistically need it to know or reason about. In fact, I'd say we reached that point about 1.5-2 years ago. There are two other barriers that remain unsolved:
1. They're less dependable than humans and can't be meaningfully punished or forced to make up for mistakes, so you can't really replace humans with them, not without having a human babysit.
2. Most people don't really have a special need for an LLM in their life. They may like that it answers questions or helps you polish a resume or, I guess the labs' favorite, helps you make restaurant reservations. But this sure isn't worth $200/mo for most people. Probably not even worth $5/mo.
It'd be kinda funny if we create superhuman AGI and then no one has any real use for it, perhaps except for military murder-bots. There's always market for that.
Even ones that cannot be easily dismissed (like the agent at the top of Google search) are often subtly wrong in a way that requires significantly more effort on my end to parse the loquacious output to determine where the inconsistencies are.
Look, can they be useful to generate the html or jscript for humanity’s 9 billionth iteration of a web form? Sure. But man, I wish sanity had reigned and people had applied ML models in general to more substantive and beneficial projects.
(Before anyone chimes in, I’m familiar with implementations of ML applied to esoteric domains; but by their very nature these don’t get all the news cycles, or hiring, or any of the other absolute insanity that the domain seems to contain)
I think LLMs are valuable and spend most of my professional life working with them.
I do wonder though whether there’s a Ponzi scheme aspect here where as long as the “frontier” can keep outrunning human review and comprehension, LLMs are always going to looked way more valuable than they are and the bubble will continue.
This started with deep learning, expectations weren’t met and people started looking for value, then GPT came out and people got wooed again and forgot, then coding, then math, cyber, etc. As long as the dust doesn’t settle we never have to reflect on all the shortcomings and can just stare mesmerized at demos.
E.g. some prominent folks definitely think LLM-based approaches will hit a wall, perhaps LeCun most notably. I also read a guest post on Terry Tao's site (which I really liked) that argues that, for all the very impressive recent AI results, they still operate within the "convex hull" of their training data: https://terrytao.wordpress.com/2026/09/13/happy-those-able-t...
I'm just curious if there is any actual data or evidence that takeoff (i.e. RSI, "the singularity", whatever you want to call it) is inevitable with current approaches.
> Somewhere around 20M people (0.2%) see first-hand that large, complex projects that used to take them weeks/months can now be completed by agents with a prompt.
No, they can't, at least not any semblance of quality. The cases we're seeing where this does kinda work is in ports and translation where all the rules are already documented in the best specification language possible with a way for the LLM to verify itself: code. We saw this close to a year ago now with Cloudflare and NextJS
> The impact scales with ambition, problem size, and horizon. A question with a paragraph answer barely stresses the system. You need a reservoir of big, difficult problems that you really care about
These are operating on different capabilities - AI's ability to answer informational queries as a chatbot frankly sucks and can't be trusted without verifying it. I run up against this every day. A problem with a verifiable answer on the other hand it's very good at solving. He knows this (his next paragraph) but he's putting them on the same scale of "stressing the system" to attempt to add proof to his introductory claim
And then there's the completely unverifiable scare that there are internal frontier models way beyond anything we've seen "swarms of thousands of agents collaborating over weeks on software mega projects: minting zero days, running cyber attacks and defenses at machine speeds, discovering new science, advancing the frontier of mathematics"
I dunno. I haven't been able to set up OpenAIs remote codex connection, their shit is buggy as hell and the web UI keeps crashing and making messages disappear. Is this what their internal superhuman "Things that would have taken top professionals in the industry years of work" looks like? Granted Claude has been really smooth, but still..
Amid all the discussion of sigmoid curves, and where the "LLM wall" will materialise, I think few people would have predicted that the real wall in LLMs would end up being humans' capacity to verify the output.
What I fear is that people simply eschew human review altogether, considering we're talking about the industry that came up with the "move fast and break things" credo. Human review of LLM-produced code where I work is already a farce, and we're not special enough to be one of Karpathy's 5,000. I do my best to manually review anything that's my responsibility, but I'm literally one of very few people left working on my team, so in practice what happens is I submit PRs that are at best glossed over by completely unrelated teams for security, malware/prompt injection, and other serious concerns. Quality insofar as vetting others' code has completely gone out the window and it shows in the number of bug reports that come back, often themselves written in Claudease. This is all on top of everyone cynically phoning it in in the first place, due to the omnipresent sword of Damocles that is additional AI-driven layoffs.
Worse yet all the incentives point to this being the most economically viable thing individual companies can do. I think it goes without saying some type of regulation here is urgently needed, and that an unexpected cause of an AI bubble pop may end up being that humans simply aren't able to keep up with the pace of the output - leading either to precautionary plateauing of capability, or major liability risks related to a decline in quality.
That was the dominant concern in the circles I’m in, so it’s worrisome it’s being treated as rare.
The response of many mathematicians to the recent dump is a good example: verifying these proofs amounts to unpaid labour for OpenAI and wastes time that could be spent doing publishable work which ultimately results in money or personal success. The slop factor also compounds the work required to verify the output considerably.
For mathematicians, programmers or anyone, if the work required to deal with slop passes the limit, it is no longer in their own self interest to use LLMs. The expectation that people will use LLMs for the betterment of humanity against their own financial interest is baffling.
Is it effective? I'm not sure, cause if it was I don't see why OpenAI & co are not releasing such a project instead of demos.
That clearly didn't happen :)
> In the next decade, they will do assembly-line work and maybe even become companions. And in the decades after that, they will do almost everything, including making new scientific discoveries that will expand our concept of “everything.”
The future is already here. It's just not very evenly distributed.
Caveat: that same tiny group is employed by the AI vendors, meaning it's in their financial best interest to make it sound like the curve is going vertical.
Considering the investment in AI, the lack of moat, and the increased inability of any of the big players to come even close to profitability (with OpenAI already breaking the "ads" emergency glass option)... perhaps he meant a tiny group is already seeing the line crater
Do you still write code by hand?
AI models can dismantle billion dollar industries. They can reverse engineer Adobe and Microsoft products that once were their moats and titans of their industry.
Why do you think they'd need to lie?
LLMs are great and will get better but we have not seen an open source office suite built by LLMs come online yet, and even if there was one I'm pretty sure few would adopt it, cause that's not the only moat, LibreOffice has been around for decades.
Sure they didn't build it from scratch but what I'm saying is that there's no moats anymore. If you build something, they will tear it down to its composite pieces and assimilate it.
There is in fact no indication of this, not evem the precious benchmaxxed benchmarks ya'll love to reference.
There is however a exponential curve of slop, and an ever increasing number of people who's minds are completely captured by these things.
I'm not arguing that they cant write code, or write a proof. Its just not written or designed well and is absolutely brutal and soul crushing to work with. Look at these proofs they're producing also, they're millions of lines of Lean that are impossible to reason about.
The way we're using the term 'verticle' to describe a curve means we're not being honest about this. This curve can actually be plotted, you can go look at the curve. It is not in fact 'verticle'. Each model release is climbing single digits on benchmarks it was overfit for.
Let's say, for the sake of argument, that the models are some multiplicative factor better on the inside.
Doesn't that mean the demos should work?
agi-society.org
Also, yes there are still bugs in Claude Code. I experience them nearly everyday.
It is markedly better than early days, but still not the best harness.
The best software written with agents seems to come from people outside of the labs (see pi, for instance — or all of cloudflare’s recent work)
Which makes me question either the model, or the holders …
Also, not mentioned in my post:
- Mitchell Hashimoto
- Prime Intellect (and all their agent experiments)
- Geoff Huntley (see Jiti, for instance)
(many more)
There's a ton of interesting software being developed with these models by people outside the in group, but I find most of the software from these big guys to be ... bland. Buggy copies of copies.
I have a version of "Cloudflare OS" running in production, using Temporal as the orchestration and MicroVMs for the sandboxing.
Sorry, I get your point but personally I don't get the hype around them.
Thanks, but no thanks, I don't need that kind of code in anything I'm working with.
Claude Code still can't get reasonable performance without writing a "game renderer" (that doesn't work)
Claude Code is still written in JavaScript, eating hundreds of megabytes to make a shitty TUI whose literal sole role is to send API calls.
Claude Code is software made by amateurs.