- Write good specs into files, usually Markdown.
- Go through the spec files with the agents, have it point out holes and problems, and then update the specs. Go through a few iterations of this and then start your agents on development.
Garbage in, garbage out.
When I'm building something new there's a good chance I don't know about a lot of problems (unknown unknowns). How do you overcome this with just spec files?
I keep coming back to this and the only way to develop properly without tons of hacks patching bugs afterwards is by working at a low level. I don't really write code but I work at the code level still and have Claude write each function and whatnot.
The problem we are trying to solve was never to write code, was to solve business problems
"I will never need to understand the code again" is a much different statement.
If you don't understand the code, and the code wasn't written by any human, when you're eventually painted into a corner, how hard is it going to be to get out? Will it be easier or harder than maintaining an understanding through however long it takes to get to that point?
You may be betting your company on the answer. How sure are you?
People come and go, documentation isn't great.
It's no different.
If I need to understand complex codebase with thousands of humans touching nowadays I am asking AI anyways
For code I care about, I audit every single hunk as its produced. I give it extensive style guidelines, and crack down on things like a 20-line essay in a comment.
For mission critical code, I write the code myself and have an agent review it.
I'm still on the fence for tests: I don't really like to write them, but the LLM-generated tests are pretty bad, even the frontier models on xhigh thinking. I usually generate them, but I don't really have the confidence they test anything except 1==1. Unfortunately it's hard to justify the time spent on writing them manually.
It's heavily customized, replaced most internal systems via plugins, a set of custom agent instead of builtin ones, different model families for different sub tasks. (don't have claude review its own code)
I'm working on polishing them up and porting a few more from my own harness, then will be open sourcing. Keep your eye out for a "better-opencode" plugin suite, I'll be sure to share it with HN :]
In the near-term, I really like GLM 5.3 prose for code explore / review, give it a shot, it's cheaper (flash model) and catches all sorts of mistakes from the Big Ai models. We fully rolled out our custom pr-review on glm-5.3-flash last week. Most devs are still on claude, moving them towards fireworks and opencode.
OpenCode Go is a great way to try out open weight models for $10/month