Have to disagree with this line of thinking, and I think it's even more dangerous with intimation further down that using an LLM for coding is akin to being a manager of people.
You do not have to consider the feelings of an LLM. You can micromanage it as much as you like, and you can modify it's thought processes. Don't like this one? Switch to another model. Most importantly, you cannot /new on a person and get a fresh state.
Working with people means building trust, respecting their autonomy, and considering how your actions affect them. I worry a lot that the transactional nature we treat LLMs is leaking into how we treat other people. And yeah, there's an argument there about how that was already the case, company / society culture and so on, but I mean additive or multiplicative to that.
LLMs feel to me like an entirely new kind of interaction, and I'd rather it was discussed in ways that set it away from how we interact with people, making the differences clear.
Yes! This is one of the major challenges I still have. It's easy to say "fix model errors in AGENTS.md / the system prompt." But finding a way to prove the effect of that change is the hard part.
It's so frustrating when a model says "no, your guidance was fine, I just didn't follow it, I'll do better next time..." It just makes me want to yell "No! You won't do better next time! This is the problem!"
> In that respect, working with an LLM is much more like working with another person than working with traditional software.
Yes, but if you work with another person, at least they'll have a chance of remembering suggestions you make.
The simple answer is "Have the AI do a refinement -> eval -> refinement loop." That pushes the problem somewhere else, into the "what goes in the eval" question ("and make sure you don't overfit" -- something AI-driven prompt refinements are not good at).
As soon as you add LLM fuzziness to a codebase, it seems to propagate like a virus. You can’t get anything reliably useful for machines downstream of a prompt—even with structured output & friends there’s always the possibility it’ll be flat-out wrong. Reminds me of the old “function coloring” issue except it’s probabilistic instead of async. Obviously it opens up incredible possibilities but I do miss the days where you could predict exactly what would happen by reading code.