They have tried to find the highest that the market pays for sota models; however, on the consumer side, this is just too confusing and unsettling:
"You can only use Fable for a week as a part of your plan" "Be ready! You have to start paying per token!" "Nevermind! we extended it for a couple more weeks" "Wait, now it's up to half your usage" "Ok, now its..."
Most people want to not care. We want our AI like electricity -- Kind of just there no matter how easy/hard is for the supply. You don't want your electricity company to be on the brink of cutting you off any second.
That's Anthropic. You don't feel they want to give you a dependable service for an, albeit premium, price. It's a constant bargaining game. That forces people to look beyond the walled garden. There, they find models that are fine... and without the shenanigans.
It also didn't help that the government yanked it which adds another source of anxiety since OpenAI is on much better terms with the administration and the administration seems corrupt enough that they would mess with Anthropic if they got a big enough donation from OpenAI.
But anyway after Sol entered the picture, I don't think Anthropic can get away with this as much and I also think they're going to face a massive backlash from Max subscribers if they do end up ending the +50% promotion at the end of the month because Sol is a Fable peer and priced very competitively.
But they're getting killed on token cost. They have to get people paying more for tokens. So then they put Fable in the $200 plan and release Opus 5. I'm suspicious of Opus 5. It is mostly worse than 4.8. It _seems_ like they nerfed it to create more distance between it and Fable.
So what have most of us done? Stayed on Opus 4.8. The statistics bear this out. 4.8 still dominates.
Now they're stuck. If they take 4.8 away, everyone will riot. If they make Opus 5.x better than 4.8, they disincentivize everyone from moving to Fable and most importantly, paying more.
Really, all they can do is take the L for now and just let 4.8 be the apex of the $20 pro plan for the foreseeable future while they work like hell to make Fable THAT much better that it earns the $200 to $infinity that they really want everyone to pay.
I just kept a $20 plan going for use on my phone.
I don’t doubt people are hitting it… shrugs
Nowadays Codex handles the bulk of the implementation and Fable/Opus on the planning.
Not sure if Anthropic patched it, but early on its release the web UI Fable guardrails will trip if you mention you're a biologist.
2nd time was literally me being lazy and telling it to commit and open a PR.
Very strange.
I've honestly stopped using CC and moved to codex. Sol has it's warts but I've never once hit limits on a 100€ and I get a similar level of performance for what I'm doing.
I wouldn't mind bumping Fable to 200$ plan if it was actually better but between the insane caps, reverting to opus/sonnet randomly and having similar perf as OAI - I'm done with it.
Next step is to put 100$ into open router and try some western hosted open models with Pi when OAI starts pulling up prices.
Reverting is annoying but being flagged for security questions while I'm doing code review is insane.
You want to disable "Switch models when a message is flagged"[1]
[1] https://support.claude.com/en/articles/16049681-why-claude-s...
What I don't see is vast areas of industry finding $10s to $100s of billions of value in LLMs. There's no lint or compiler that can check for correctly constructed contracts. So LLMs, which should be useful to law firms, incur a lot more manual checking of their work than coding agents.
Less formal document production in other industries is likely to have less structure. That might not matter in some settings but I'm having trouble thinking of an example off the top of my head.
I don't want Shakespeare, I want Bob the builder.
Half of my work is telling claude how to behave. I'm pretty certain they have enough _data_ to realize people do the same thing time and time again.
Check this comment of mine for a better explanation of this: https://news.ycombinator.com/item?id=49413353
But people are not as alike as you think. I doubt I share your unique preferences.
That said, I don't spend much time telling it how to behave. Are you sure you're not fighting the default system prompt?
Either way, a model used to solve the top 10% of problems that people use AI to solve for, being used 10% of the time … seems like it’s in a decent place.
I find Fable indispensable, and measurably better than alternative models, for complex feature development in an existing codebase. It’s the closest thing I’ve seen to nearly one-shotting features. Even still, I only use it for the hardest features and Opus 5 does a good enough job on the rest.
It doesn't mean hosted frontier models wont exist, they'll just be rare. It's no different than any other commodity market, for example most cars are cheap commodity models, with rare individuals buying expensive luxury cars and businesses buying expensive trucks and specialized equipment.
I can hand hold opus but I would rather just ask fable to do it and give me the result that I review and works. Opus will waste tokens and still require me to help nudge it in the right directions.
I think the next gen models from china will put us in a spot that the cost can plummet and I won’t need the Sota from anthropic
Also, Sol doesn't refuse constantly and speaks like an engineer rather than a deranged academic.
Opus 5 is legitimately terrible and can't or won't follow instructions. It is of negative utility and does more harm than good to my codebase.
Most of the benchmarks have exceeded their usefulness. Opus 5 beats fable 5 on many of them. Anyone who has used both models will notice immediately that this doesn't translate to the real world. Opus 5 is nothing short of a regression from Opus 4.8. Fable is genuinely a great model so long as you don't trigger a guard rail and it downgrades.
Sol in my experience isn't significantly different than fable ignoring that Sol burns usage 10x faster but the end result is hard to differentiate.
GLM 5.3 is a hair behind these two.
An anecdote but not an original one from the people I talk to.
But Fable security false positives and pricing just make it not worth compared to Sol IMO.
I would compare it to a extremely high end $15k PC, or an expensive pro-grade video camera, or a freight train, or a …
I would say at least 95% of the global population will not encounter a situation once in their life where it would be actually useful/warranted.
It’s more reliable and makes less dumb errors than Opus.
It still messes up, of course. But for my working style, I definitely prefer it.
That said, Opus 5 is broken. Use 4.8 or another vendor for the build agent.
Yes but those issues will be much less severe with Fable-written plans than those written by lesser models. I know this because my workflows at both my regular job and my startup involve multi-step agent reviews via codified adversarial review skills. Fable as a reviewer will frequently find blocker-level issues with plans written by GPT 5.6 Sol, and sometimes with Opus 5. The opposite almost never happens. In fact I cannot remember the last time it happened.
Cheaper, more powerful AI will continue to expand the bubble. Projects will get more ambitious. Everyone will build out their own custom little software. Code diversity expands and requires even more AI.
They are already good enough at what they mechanically are.
You have to use the right harness, right verifiers (automatic where possible, human where not), etc much much more specific than a generic one like claude code or codex, and it will also be able to work within constraints and be the "proposer" of an imaginary optimisation problem and an excellent one at that. But you have to frame the task at hand in that manner or maybe even reorganise the task you do itself so it is more amenable to being framed that way. If you use it this way, it is _already_ massively economically useful. But it will take many years for it to actually be usable in that way, since you need DC capacity to come up first which is few years away and also well, massive organisations that have to integrate these will usually take many years to do so.
It is also useful albeit less so in cases like general SWE, where you still need a human in a loop for non-verifiable requirements, and also in other general usecases where information retrieval is too intractable and you need to carefully use LLMs as a component of the overall system.
I am not saying Fable or whatever the biggest models are are useless - they will certainly be useful for tasks at the frontier of the day - which is today complex exploits and open math problems, and well, tomorrow it could be something in biotech. But this is not what the entire bet is on at all - just automating day to day drudge at the tens of thousands of massive companies and governments we all know and love is more than enough. With the right training data (which _also_ is a bottleneck and takes time) you could even automate certain processes entirely. Sure, if we get a crazy medical innovation and end up saving trillions in healthcare great, but that's just a bonus.
None of this is to say that I think there is zero sketchy financial engineering going on
And you can have extremely simple work that you nevertheless have to do as a professional. For example, drafting a routine customer email, summarizing a meeting, formatting a report, filling in standard documentation, or making a trivial code change.
So I don't think the $15k workstation / professional camera analogy really holds. Those are specialized tools whose capabilities are mostly useful within a fairly narrow domain. A general-purpose AI can be useful across thousands of completely unrelated tasks, including one-off problems encountered by ordinary people.
What if:
- Opus 4.6 was the last Opus generation that got a lot of use by Anthropic's own employees
- After that they primarily used Mythos internally
- 4.7, 4.8 and 5 were RLAIFd by Mythos "teachers"
- Hence why 4.6 is the last Opus gen who doesn't report back like a robot wanting to cover every potential hole another AI system would've spotted and criticized
- Hence why coding style in Opus 5 also gets criticized, not only behavior in CC
But a while ago, I had access to a Fable harness that just never gives up. And that verifies itself. It burned $100 in API tokens in 15 minutes ... but it succeeded for all the prompts where Fable + Claude had failed.
And I believe that's a real issue for Anthropic. Fable+Claude is not too expensive thanks to the subscription, but Claude severely nerfs Fable. To save money, I guess. Fable API + Custom Harness is a different class, it's so much better. But API tokens are so expensive, you're cheaper off hiring a freelancer.
This is the first time I've felt, and I use the word *felt* since I don't have a suite of benchmarks or any sort of material approach towards comparing models, that Opus has declined in quality compared to before. Primarily I think its powers of deduction and understanding, even on xhigh, have become much worse. Before, being vague and providing a simple prompt would be enough, it could deduce and expand the details it needed, plus ask you clarifying questions, now this is no longer the case. A concrete, personal example, for a personal project, I've asked it to setup ssl over local IP. I didn't go into too much detail in the prompt as there are many approaches it could take and I didn't care too much to choose. It did horrible. The first thing it did was say the best lightweight approach is to add a reverse proxy. I'm like ok, makes sense. Then after asking it to proceed, it went and added a bunch of config to my golang service and didn't even setup a reverse proxy even when it said that is the way to go. It even said it didn't set it up lol. Then after I told it to do so it failed building the config in a way it was asked of it (support LAN IP and tailscale IP). Etc etc...
When Fable came out it was huge, the benchmarks told the story, and the story mostly matched the experience. It felt, again, intentionally saying felt, like it was miles ahead. Now benchmarks say that there are many models that are close, but in actual use Fable still *feels* much better. I think benchmaxxing the new open weights models is ruining the value of benchmarks, if they ever had any. When you actually put them to the test you see 500k tokens of reasoning with "Actually..." and "Wait..." in every third paragraph of their reasoning trace.
The price for Fable is definitely too much for any personal use now that it's no longer included in the subscription, and GLM 5.2, Deepseek Flash and Qwen 3.8 served locally or via cloud provide a lot, requiring a bit more babysitting though. Considering the price of Fable, my 5k USD Epyc server would pay itself off in less than a year if I used Fable or Opus in the same manner so at least for me the decision seems easy. And considering the point I'm poorly trying to make, that Opus doesn't feel like frontier anymore, this is probably the last month of my Claude subscription.
I switched from Opus 4.7/4.8 to test Kimi K3 a few weeks back and the test hasn't finished; it's my daily driver now.
Given their general behavior and preference toward social engineering to scare the shit out of normal people...this seems fitting.
if people dont see why they need a model this smart, they probably arent using ai enough
I am not doing awfully complex tasks though. I imagine a lot of other people are in a similar boat, either switching from Claude to ChatGPT or even just min-maxing DeepSeek V4 Flash 0731 or similar.
Anthropic in particular is much more compute-constrained than OpenAI and SpaceXAI and has relied on partnerships to provide inference. This reality factors into their pricing and usage limits (they started 'adjusting' the 5-hour limits during peak hours, and it certainly wasn't an upward adjustment). Accordingly, this is presumably what Anthropic wants, given they develop and release the lower-end models, suggest users use them in various nudges within their product, position the bigger/more expensive models as "For the most complex tasks" in their UIs, and so on.
I think the problem (for Ant/OAI) is that there is no sensible lockin or moat. LLMs are essentially interchangeable and stuff like a harness doesn't offer enough value on its own for someone to be locked into using one of them.
Now with the onslaught of the Chinese models that offer almost the same quality for much less money they have a very serious problem on how to proceed. Investors now might be looking through rose tinted glasses but their patience has its limits.
Enterprise is a whole different ballgame IMO with countless technical, compliance, and employee adoption considerations that add friction to switching. It's also where both Anthropic and (as of last week) OpenAI get the bulk of their revenue, and, incidentally, the venue where US Government regulations on Chinese models would have the most impact.
Nobody has complained and seems like for every use case we have Opus is more than powerful enough, especially with Opus 5
I find it funny how OpenAI got caught lacking for a very brief window, but it turned out to be a very critical turning point.
Like a guy that that's at the top of their game the entire year, and the one day they have the flu, the CEO does a surprise performance review.
Sure there are Microslop and Oracle db users but most of the world we live in is Postgres and Linux. That's why I think most companies will run llm's like that.
> good at its job just use it instead of re-inventing the wheel
Exactly why is everyone reinventing a harness every month. There will be Microslop / Oracle harnesses and there will be 1 or 2 open source ones that win.
Most of the people pushing this are just hoping that they can cash out before the hype pops and financial gravity crashes the party. Sam Altman recently claiming that the singularity is here is so stupid on its face he should just be treated as what he is, a huckster.
None of this stuff ever made any sense on what it was being sold initially. It was always insulting that the media and business leaders tried to argue that the tech could replace entire call centers or vast swaths of entire industries.
People keep arguing, but it will or it has based on extrapolating certain, reasonable use cases. Klarna has shut up about replacing call centers with bots because Markov chains with memory only can do so much.
I’m happy with Claude. If they become (bigger) jerks, I’ll switch to something else. I don’t ha e the energy to praise Anthropic today and I won’t have the energy to demonize them tomorrow. The emotional investment people have for/against these companies feels like celebrating or being offended by the weather.
TBH I don't think any of that is unique to the IT industry's relationship with the Valley. Other technology-driven industries have a similar worship-ish relationship with a few rarified businesses. But the culture of the IPO exit accelerates all the most short-term motivations to do anything.
No shit we're all paying 清冲 Flash to do the grunt work. Turns out though, paying 清冲 Flash a few more cents does exactly what Ivy Wasp Pro Mythical does. Crazy how that works.
I can literally open a new chat with just "Hello" and it gets bumped.
Claude keeps track of memory if you've turned that on so all conversations are somehow tracked over time Claude randomly will mention that I'm a developer while I'm asking unrelated questions and say oh because you're a developer you might like this or because of my background and infrastructure you might find this interesting and I always get creeped out by it.
They have the concept of incognito chats, but I can see those being worse without context from other conversations happening.
Regardless, the number of innocent biochemists who have been hamstrung by Fable is essentially zero (but somehow they have all shown up on Hacker News to decry Fable), and the number of non-innocent would-be-biochemists who would love to have unrestricted Fable help them with their non-innocent biochemical research is absolutely greater than zero.
In my view, from the testing I did since Thursday, it's better than Fable. I had just finished a rather large task that Fable completed, including a /review and an /ultrareview.
Ox Alpha found bugs that Fable and Opus missed, and it continued to build things like a pro.
It does have issues with availability - but it's on a free promo right now. That also means that I don't know how much it would have cost if I had to pay API prices for it, which might not be cheaper than the subsidized small company / consumer usage, but for large companies paying API prices for either offering, the difference will be considerable.