I started to gloss over hard after a few paragraphs because I don't really have any the problems you describe anymor; at least not to the point i'd be blogging about them. Instead, I've just been iterating with an agent on various pi extensions that solve the issues.
As with operating systems - you can hold out in the hope somebody solves the right mix of issues in general, and they match your situation well-enough.
That's your choice.
I would note though that being the passive consumer gets you either Mac or Windows UX & prices.
Isn't this the point of Pi to implement all the features you need yourself? Complaining about lack of features in Pi doesn't make sense to me, it's their whole identity.
Although I must say that describing specific problems is valuable on its own. It's just (a) why stop there and (b) if stop there, why shape it as a complaint and not as a list of ideas for others to build.
A lot can be improved, but this is already so much speed.
I still occasionally have issues with open-weight models, but the frontier labs have solved the above for most use cases.
This is a pretty big failing, which is compounded by the fact that most humans don't know which model to pick, or make assumptions based on Anthropic's hierarchy or "effort" involved.
Like Fable: your toughest challenges. You mean, like Fields Medal toughest challenges? Or analyzing and updating three monster spreadsheet toughest challenges? Or writing a new novel in the style of William Gibson toughest challenges?
The descriptions are near-useless and tend to flip around, as model families are not released in sync anymore, that's true, but fortunately, thanks in a big way to subscription pricing, the choice is simple: start with the best model on offer, and when you run out of quota, downgrade to the next best (or briefly switch providers).
Codex Astra can do a great-(ish) job as a project coordinator dispatching tasks to a pool of 6.1 Sol sub agents. You can even give it an explicit goal and ownership over ensuring the work is carried out efficiently.
However the OOTB harness(and prompt) configuration may not do this for you. You'll have to provide guidance over how you want it to operate through your prompt, a skill, or etc.
And I'll say even though it's really good at this.. Having even more layers than 2 can help; a single agent given too many responsibilities will start to become fixated on a number of them while neglected others. You can check in occasionally to "nudge" it or you might need to split out responsibilities more..
I will say it's crazy Codex doesn't have more built-in task and sub agent management features. I almost wish that it had some stock orchestration patterns that worked OOTB, and then you could opt-in to a leaner setup where you provide more of the instruction.
Makes me wonder about doing some archaeology and trying out really old harnesses on modern models...
I let most agents work asynchronously and don't pay attention so I don't care that much about the sequential nature. But if it's a problem for you then fix your harness. This is a bit like saying "Why are shoes so shit? There's a stone in one and it just gets stuck there and your foot steps on it and it hurts". Take off the shoe, and shake out the rock. Put the shoe back on. You have the power.
Edit: I have access to codex, vscode, GitHub co pilot cli and all anthropic and openai models (excluding mythos).
But maybe that’s on us, AI doesn’t care about all these special cases, it’s not debt to it as it will simply read them all when making changes. We’re obsessed with quality and what code is supposed to look like but those are human standards, AIs evolve to look at this complexity as a single picture, they can simply see through it so what is spaghetti code to us is merely some code to them that works as it should and is efficient. It’s interesting we can see how the two things drift apart, you would think at some point AI generated code should explode but it hold together unreasonably well in most cases…
They even have split my decisions to human decisions. Proposed and approved work. They can iterate on approved work without me just fine.
And I only just started with agentic coding in last few weeks before that I was mostly a copy paste chat person.
Your shoe analogy also breaks down because really the shoe is the issue, not the stone. And expecting everybody to make their own shoes is, well, I mean we just don't do it that way anymore for good reason. Let the cobblers make the shoes, and the runners wear them.
I run an LLM server with Qwen 3.6 in the office, and OpenCode, which the OP mentioned, usually defaults to sequential TODO lists, and it works fine with our little LLM server with 3-4 parallel users. But I noticed that once in a while the LLM got overloaded with requests in the queue, and you couldn't do anything for 20-30 minutes. My investigation led me to an employee who used QwenCode. I tried it myself then, and indeed, it immediately launched something like 6 parallel subagents, where OpenCode would have sequential TODOs with the same model by default.
So in the end, I had to detect QwenCode on the server side and serialize all its parallel requests into a single request queue, because it made life miserable for other OpenCode users :)
One thing that seems to help for me is to do the docs before plans (collaboratively edit with agent). Then I understand what this change is going to look like from the user's perspective before we start implementation. This seems to help keep things on track.
While I don't use this plugin a lot anymore, I think doc-driven development is one of the most effective ways to do development in any paradigm, I should probably refresh this plugin and use it more:
https://github.com/tmpdir-org/tmpdir-claude-code-marketplace...
Have you tried telling models about your dream agent environment?
They can build it.
This is a great list for future Agent / Harness software engineers to read!
Oh sure, there may be some Agents/Harnesses that already accomplish some of these things -- but there doesn't seem to be one (as of the present day that I write this) that accomplish all of them...
As someone that watches the Agent/AI Harness (and related software) space, I will definitely be referring back to, and re-reading this list in the future!
An excellent post!
I have no idea what stops that person from just making it, with an agent of course.
I asked DeepSeek to translate a page to five languages and it opened five subagents each one working independently on the translation, once they all finished the main agent informed me of the job completion with a bell. Fantastic!
Sooo, which agent?
As expected. A model's knowledge is what was it ingested a creation.t
Unfortunately what we get is worse - for the same reason. Model version thinks it is its previous version.