I also changed the way I do review because it has been more than a year and the juniors/new hire are still lost, wether on domain knowledge for the older new hire, or just capabilities for the juniors, and discussing with other departments, it's the same for like 95% of them. Now, rather than correcting the PR or adding a request for change, I add a whole unit/functional test to the PR and let that as an exercise to pass the test. They can use AI but I tell them to try to find what part of the code doesn't work before generating the fix, hopefully they'll take ownership of the code if I keep doing that.
this makes me kinda sad. they don't do this on their own? are they not even a little curious
So, the models ARE dumb. They just are very good at finding patterns in their training data. I mean, when they code minecraft clones, it is not because they can cook up how to write minecraft clones, rather, their training data includes a lot of minecraft-like games code, and they just reuse that.
Don't get me wrong. It's not great. It would never pass any of our policies for things that actually operate stuff on the power grid, but as an administrative tool that can live in total isolation from the vital networks. It's perfectly fine. It's also not like we would have hired the best software companies to build it otherwise. We'd hire some low-level cheap consultant house who would then likely get cheap student labour to build it. With that in mind though, the AI is much better than what the realistic alternative would be.
Money wise it's also cheaper. It's been roughly €1000 + the time it's taken us both. If I had known they were doing it, I would have rolled out the developer cowork app/skills/whateveryoucallconfigurationsthesedays to them. This would have avoided their AI building it to be depoyed on a VM rather than in our managed k8s in our Azure. It would also have written the code a little different, used UV and maybe django rather than flask. But hey. For what it is, it's like a 90% cost saving compared to buying what would've been a less maintainable and lower quality system.
I think perhaps the greater issue will be finding people who want to extract the gold from the heap of shit and getting it to run in production. I don't personally mind, but it's not like any of my colleagues would've wanted the task.
The insecurity in a vibe-coded web portal isn't that someone hacks it with XSS, it's that after the next vibe-coded release, some X quietly becomes −Y somewhere no one expects.
From this perspective, having no software at all might be better, or as in your case, safer.
As you point out this portal isn't that, but what protects us is the processes around compliance. This can't grow from X to Y because not even the CEO has the authority to overwrite our compliance gates. The EU is a tremendous help in this area since personal liability changed things completely.
There's not a lot of software where users dont really care if it goes wrong.
However, over time the complexity of problems where you can get away with less/no oversight is increasing. And the models are already great at solving certain classes of problems where one doesn't really care that much about code quality, that wouldn't have even been attempted in a pre-LLM world. Over the weekend I was using Claude to add features to the compiled (no source available) firmware of one of my audio devices, adding workflow features by patching assembly and custom DSP code.
In coding, as with other areas, what's emerging is jagged intelligence.
Why not use actual frontier models, and you know do some real research, before writing a blog post?
My point is, there is a monumental difference between flying Ryanair and Emirates First Class for example.
It is funny that only real research on productivity gains from AI shows at best very minimal gains, but AI bros will always tell you "no no no, you have used wrong model, try a different one, there are more of them, you have to try, trust me" and call that a "research".
Looking forward to the authors next blog post on how after driving one car they find that all cars are slow and uncomfortable. Then the follow up after trying one phone and that all phones have bad cameras and battery life.
Wow. I guess that’s the punchline!
Actually, I think the author put DeepSeek on purpose to avoid the obvious ChatGPT/Claude comparison — because whatever he chose, there would be a question of why model A and not B, while the point of the article isn't about models comparison at all.
Imagine someone trying to make the case that riding bicycles is a terrible experience and their whole argument is that they took a random cheapo bike with flat tires and rode it for 3min and that wasn't fun. Sure, but if you buy a 25k carbon bike you will have a different experience. I'd not trust that person. If someone told me they have 10 bikes they ride daily and can explain the differences, in detail, between their bikes, and what they excel at. I'd trust that person's opinion.
I'm also convinced the effect would not be there If I had a 10$ budget.
Sure it would be. I pay $10/month to OpenCode for a Go subscription, it's fine for day-to-day coding tasks. I wouldn't necessarily try and one-shot a production app on that budget, but with decent planning and test-driven-development, it gets the job done
Actually, I lied, Codex is free and I developed my first AI written application using it and the free tier limits were generous enough to work on it for several months.
The author makes a good point. If you don't know what you're doing, AI accelerates that. No question.
But they put a whopping... ten bucks into using DeepSeek and weren't impressed with the initial results.
I know they try to cover this with "you just aren't prompting correctly!" but if, in 2026, you aren't able to have an LLM generate decent quality code... IDK what to tell you. Good luck I guess?
> people are doing real paid work using LLMs that they could not otherwise do
FWIW, this is not exactly new; those same people were just using other sources like Stack Overflow, blog posts, etc. before, cobbling together random code snippets, libraries, and so on without actually understanding any of that at a relevant detail level.Sure, with LLMs, one can naturally tailor this much closer to the current need (or at least the need one thinks they have) and iterate ("spew") faster, but it's not a new phenomenon in general.
Any significant testing will inevitably create that situation. The LLM won't have enough context to handle the more precise business requirements. The dev will have to read the code carefully and make their changes by hand. Additional rounds of testing may cause thrashing between regressed states until something clicks for the developer. That lightbulb going off is called "learning" and they are human after all!
AI is a godsend to this cohort of code monkies.
Edit: we graduated in 2006.
i know of several engineers who produce absolute slop and who probably would have produced nothing at all in pre AI times (which would have been preferable) and probably let go or never hired (even better).
they impose such an enormous drag on productivity that they more than wipe out any gains from people using the tools responsibly.
Now with AI I can 10x my output and 10x my bad code!
Airbnb, Stripe and Dropbox were created in a different time when the market was much less competitive.
Saturation of software development velocity doesn’t increase large scale product opportunities in the market. It can also mean that opportunities get filled even more quickly by niche players, and nobody gets to grow to Airbnb scale.
IMO the latter is what’s currently happening. AI-powered companies are like little mammals scurrying around between the feet of the dinosaurs, and commentators like the OP look at the evolution of the brontosaurus as evidence that the mammals don’t seem to be growing as they should.
You won’t get capital let alone VC if you’re not AI.
It’s infected everything much like crypto did just 3-4 years ago.
That's one of the few consumer-facing areas where there are still standards in place, namely PCI-DSS. As far as I know the audits require the name of a human who is responsible for payment security. Card companies can one-hit kill your startup if you're breaking those rules (maybe purely blockchain startups are exempt).
Yes, you can offload this to stripe, but then your app should never see the card number and certainly not the CVV. You end up storing these, even by accident, both stripe and the card companies will hate you.
* stretching the timeline: the actual real programming ability appeared in LLMs in the last 6-8 months, not 3-4 years,
* using the weakest possible tool: and I bought $10 worth of DeepSeek credits that is a far cry from Claude with Fable.
Also, I know nothing about marathons but for most uses putting the app, database, and background processes on the same server is very much the right starting point. With the next steps being employing Cloudflare or similar solutions long before managing a fleet of servers.
This is fair to say if someones last experience with AI was copy-pasting code into GPT3 chat windows years ago, but Deepseek is a more than capabale model and enough for someone to get an informed opinion about the technology.
If people have actual counter argument, use those. And if some of those counter argument are "What you say isn't possible, the neweste model can do and here are examples of that", that is fine.
But a blanket "Nuh-uh, it wasn't Model X" is not only a poor argument but also automatically invalidates any criticism when a new, better model comes out - and that can't be the basis of a good argument.
The field is moving fast, and asking for scientific arguments is not realistic. It takes an extreme amount of effort to show what exactly is different.
We were in a similar position with static vs dynamic typing for decades. There is still no scientific proof that one is better than the other, but it is quite obvious to professionals which flavor works better in a given situation.
So, even though the argument might be sloppy, I subscribe to it. Using DeepSeek to dismiss better models is the bad argument here.
Edit: added "(for some)" as a disclaimer that you still need to be a fairly decent programmer to actually benefit.
You don't have to pay money to use Codex, there is a very generous free tier that costs you nothing, you just have to accept being told you're out of tokens every day. Because your token limits are low, you need to make sure that you accept or reject everything manually and when it tells you that it wants to run a command you have to paste in the command into your terminal and only paste the relevant output back otherwise it floods the context window.
And since then there's been 3 more "now real programming ability has been made available and previous stuff was just toy examples" cycles (summer 2025, winter 2025 and spring 2026)
Looking forward to the next "everything before this was trivial and bad, here's the good stuff" moment
Or B that the ideas are there but are not get released as the code is ai slop?
Disesdi Shoshana Cox
- when are all the "I am running agents 24x7 a day on max plan" people on HN gonna learn this?
- best case scenario bro: LLMs get infinitely better and nobody needs to code anymore
"software requirement prompting is the new skill" I can write those pretty well bro
- worst case scenario: AI market crash, cognitive debt spikes across every major organization filled with vibe coding juniors that have never spent a single day in their life debugging a production setup without AI.
- Now the whole world is filled with 90% programmers that cannot add 2 numbers in c++ without using GPT
- Guess what bro? I am now one of the most valuable programmers there is :)
It's like cars. If you only ever drive at 30 km/h, a 30 yo car, that had its last oil change in 2008, might be great. But, if, in the same car, you start driving at 200 km/h, you might notice that there's a lot of room for improvement. My programmer-self is the 30 yo car in this metaphor.
You can use AI in two ways, either…
1. set the plane on autopilot and arrive at the destination having no idea how you got there - like any old idiot who can type a prompt. And also pay a very expensive air fare for this super power auto pilot.
2. You can pilot the plane yourself. Plan the route in advance. Set waypoints. Make decisions along the way. And land the plane manually. And the reward is you know the route and you pay 1/10 the cost.
Neither was ever about some bottleneck on pumping out code and everything about marketing and network effects
I think this is a great question, and I have a completely unqualified theory.
I work in a fairly specialised field, I've notice when I use AI to try to generate code in fields where I'm very experienced, it produces a worse output at a slower rate than I can produce manually[0].
When I use it in fields I'm not an expert in by any stretch, like webdev, genAI massively improves my output by a huge factor.
I think this makes the productivity gains a little fuzzy. Claude code makes me 10x faster at webdev, but I'd be a very slow webdev. Ultimately, the next innovation in a field, will probably be from an expert in that field, and they therefore won't see the kind of "10x productivity boost" that non-experts do.
[0] I still find AI helpful for exploring code etc, just not so much producing output.
Seriously, that was the lowest level of skill I saw in 30 years of development.
Now we have people who can't even do that proudly showing off PRs to widely used products.
"B-b-but I do the systems design and hard thinking".
Sure, buddy.
How does a manager do this, they ask? I answer: the same way they do everything. Lunch. If the effort fails, at least they tried! Deck chairs, Titanic, etc.
[* This is not financial advice. Please don't actually do this.]
The bottom line is that markets are complicated, useful technologies are often accompanied by bubbles, investors are not always rational, and people generally try to make the best decision with the information they have available. The answer is likely somewhere in the middle, but that's a lot more boring and a lot less inflammatory.
Because going long and going short are very different. For successfully shorting something you need to have a pretty precise estimate of when the crash will happen. Being a comparatively short time off can cost you everything.
Most look forward to picking up the assets at a heavy discount. =3
It is not a question of if the bubble will go, but when... but you are right that Bears or Bulls always get it wrong predicting the future (if they are a legal investor.)
https://www.youtube.com/watch?v=wTiYaWFP59Q
Personally, Shrek movie release correlation with market corrections is funny, and a new film is due June 2027. Please hedge your bets with a diversified portfolio. =3
The "I tried it and it sucked" is borderline conspiratorial at this point. There are enough talented, thoughtful developers saying there's something real here, and it is worth believing them and investing some time to understand it, even if you come out the other side and decide you don't want to use LLMs.
Use a very good model. Set up a good harness. Spend some time on your system prompts and skills. Develop your intuitions about how the model works, what it's good at, how to scope the work, and how to steer it. Recognize when it's alleviating menial work and recognize when it's making choices you really need to understand yourself. Be patient and accept that the failures are going to be very painful for a while.
Don't write it off until you've genuinely seen the upsides.
However, your discourse is not the usual borderline religious discourse that's used by LLM advocates. You're not extreme enough either pro or against LLMs :)
As far as system prompts and skills, my approach has been to start with nothing and add very specific instructions as I notice issues or feel like I need them. Mostly around trying to keep it terse, kill the obnoxious rhetoric ("It's not X. It's Y"), and give specific guidance on how to write code. Skills are things like "use edge-tts to generate spoken audio for this answer and send it to me" kind of thing.
I tend to agree. Not with the extreme version of this statement: there are some genuinely useful instructions I give in my Claude.md and prompts. Comment style, how much to push back, which subagents to orchestrate for tests, reviews, etc.
But I've tried the single-sentence versions of those and versions with multiple files of long prompts for the orchestrator and its subagents, and I can't tell the difference in output quality. If anything, the shorter version is better
Are those "skills" or just instructions?
My definition for "skills" is the "you are the greatest software architect ever born" type bullshit.
This is ridiculous. Basically you're saying "if you try it and find it sucks, you're wrong, keep at it until you change your mind". That isn't a tool at that point, it's a religion.
Next to maintaining and expanding my own network of websites I 10xd writing boring CRUD applications for companies that were otherwise unable to afford it, making all their employees more productive. There's true economic value in that.
It's going to be a worldwide software experiment.
I don’t want to solve a new problem with an LLM programming assistant anyways - that’s the fun part of my work, why would I get rid of it?
You can most definitely ship crap much faster than you used to. It's obvious to anyone when you're shipping crap. And it seems everywhere I look people are shipping crap. Both software and writing.
The hardest part of shipping quality software is not and has never been "writing the code". The hard parts are product taste, architecture and ensuring your product actually solves the problems it should in an efficient and secure manner.
If you don't have an intuition of those, you will likely still be shipping crap. Just faster.
Where AI shines for me is accelerating the learning and exploration process. I can get up to speed with new tech fast. It is good at spotting issues in designs and code. It can knock out quick tests or benchmarks to support me. The quality of what I can produce with AI support is much higher than I could without it.
So it really just depends on what you use it for. If the goal is "replace humans and ship fast" that's one thing. If the goal is "explore the problem space in greater depth", it's another.
Tbf, I'm not sure we can really lay this one at the feet of GenAI. The SEO bros had pretty thoroughly ruined search before LLMs took off - the process just accelerated a little at that point.