Prompt was "read and update the config file with new data". This work on 4.6 takes <2 minutes to read the file, parse the new data, and patch.
Opus 5 Result: 43 minutes of pulling containers, running sandboxes, creating testing suites, which included evaluating the entire repo beyond the scope of the config file.
Both: one file modification
/on The prose is load-bearing unbearable — every sentence feels like it was engineered to sound profound rather than to be read.
Recently got approved at work for ChatGPT Pro so I could use Codex.
Blown away by the speed. It feels like using Claude Code for the first time again. I don't think Codex is doing anything revolutionary, just better handling of which requests should go to which model, and having faith in some of the "less powerful" models for more than you would think.
It seems the TUI coding experience is very much an open race. This is motivating me to look at other agents / harnesses as well (maybe Gemini, OpenCode, etc).
If I have a user input and then sanitize and inject that into a prompt to do something, I have no idea how much that is going to cost at all and no real way to measure this properly. A parallel example is digital ocean or aws, i can go and measure/limit my compute/fs/memory/startup times/etc and while it can be impossible to get down to the last flop of money allocated - i can run things on a real budget with real constraints, opposed to an LLM where I have to .. prerun a sanitized user prompt through a tokenizer and then ask an LLM to guess what it may do and give token consumption estimates and then act on those in any sane manner for the user?
Perhaps i'm missing something to do realistic and static rails on things but I don't see a serious way at scale to use the token billing model handling things requiring a users free text input short of having to go pander to VC money to throw money at it until someone else figures it out.
*to clarify my rambling... We should be billed and given controls based on resource usage itself and not an opaque token concept on top of not being able to spin any knobs that control it's resource usage.
The model providers are quite aligned with concerns like customer retention. These arguments only work if there is no competition. We exist in a marketplace of black boxes. There's not just "the one" you must suffer. You have options. You can build your own too.
Theoretically.
In reality, one sessions output tokens become the next sessions input tokens (at least if you continue the topic) so, its not as aligned as all that.
But the parent is right, when incentives are not aligned, friction will happen. Its inevitable.
LLM doesn't seem to be keen to put in effort either!
Is this AGI?
The chat-based models are obviously being lobotomized based on personal usage and general load (e.g. PST business hours are worst).
API doesn't seem to be affected by this.
I maintain a Claude subscription for Fable but seldom use it.
In that case, the agent will respond incorrectly because it has no visibility into what reasoning mode it’s in.
I was approved.
3 months later, my approval was degraded into "in review" (revoked). I'm sure my account was flagged based on contents of debugging/researching firmwares/etc.
I opened a support ticket. No response. I opened another support ticket. No response.
1-2 weeks later, I got a response that I will not be re-approved and I need to reapply. No problem.
The page to reapply on does not allow me to re-apply because it my account is stuck in an "in review" status.
https://github.com/anthropics/claude-code/issues/84352
The community thinks it's a bug. I'm 95% sure it's not and a bunch of us who were previously approved had it revoked due to flagged content and will not be reapproved.
I switched to Codex + got TAC approved instantly and have not looked back. It's a shame. That's 100% separate from whatever the heck the quality of Opus 5's outputs are. The way it talks... insane. I would bet a good amount of money their next release will focus "reduced simplified responses" if I had to guess.
$2t company by the way
Not convinced here.
The problem is that Anthropic seems to be getting away with selling one thing and delivering another. You pay for Opus, you get something else etc.