Like, even if you don't care about (or even like) his politics and can look past how unlikable he comes off as, the damage he's done to his own reputation in this domain just makes using his products like this a no-go. He's literally so rich that he can get caught personally looking through chat sessions and it wouldn't slow him down a bit. He's too rich to be held accountable, and that makes it impossible to trust his businesses. It's a funny dynamic that I don't think is appreciated enough, but I know that if Google or Amazon or OpenAI or Anthropic (etc.) got caught doing something like that, the backlash would be astounding and the reputation hit they'd take would be brutal. Here, Musk would just awkwardly come out attacking people for not letting him behave unethically even more than he already is, and that'd be it.
Beyond that, the obvious astroturfing that occurs on this site (along with reddit, etc.) when it comes to Grok isn't helping. All I hear about Claude, GPT, Gemini, etc., are how terrible they are, yet any discussion of Grok seems to always revolve around sensible, but confident, assertions that it's actually a great product and every new release is the point where Grok finally catches up.
Looking through chat histories is boring, mundane stuff. He's richer than that, think bigger. I think he could kill a random person in front of thousands, and by the next day we'd see articles arguing why the random person actually deserved it and why it's not that bad. Whatever consequences would be lined up would inevitably face unexpected roadblocks which would all result in nothing happening.
that's the hilarious paradox at the center of his antics. Musk is infamously petty and insecure. We're talking about the guy who tweaked Grok's system prompt to flatter him and paid someone to boost his fucking Diablo character for clout. I wouldn't put "looking through chat histories" past him for one second.
Ironically, I only see coments like yours regarding Grok.
Tesla self driving cars, (somewhat) as you say, but even the biggest proponents of Grok are like "oh no the best model is this, ugh".
It's less about "who is more trustworthy", it's more about "who is more willing and able to affect me".
Also, if you want true privacy you should run AI models on local hardware. (Guess which country's models dominate SOTA/near SOTA open weights? Yes, it's China, and it's not even close. You can run full-fat DeepSeek locally for (just) under $10K USD.)
It's crazy how much Chinese = bad the media or US companies have washed into you. Why lump it together?
Like any place and any company there are good and bad 1s.
It's not the Wild West over there...
It's not a matter of whether or not you can trust these governments at all; it just comes down to which government do your self-interests align with best. It's not some grand political statement to acknowledge that my interests don't align well with the interests of the Chinese government. It's just an obvious fact.
The US has much further to fall, but it's falling very, very quickly and if there's ever another Democratic president they're going to have to rebuild a lot of the government from scratch.
When the next democrat president gets into office, he or she should do the same thing as Trump: put trusted deputies in charge of various departments and whip them to actually do what people elected the administration to do. That’s how our system is supposed to work. And democratic voters would I’m sure be much happier with the party if they sometimes actually got what they voted for.
That is a conspiracy. Do you even know what happened to Jack Ma? From what you're saying you don't.
Also that was MANY years ago. The Shanghai stock market crashed. Companies had a lot of fear then yes. Things have changed and repaired. I'd say China in this sense is moving upwards and the US is going downwards in policy.
> You could argue the US has the Cloud Act
No, not really. Your Jack Ma example happened to Elon Musk to some extent. Jack Ma had a feud with the Chinese government as much as Elon had a feud with the US government in the last year or so. Back then Tesla and the other projects all tanked.
Grok 4.5 works. 4.6 is looking even better.
Grok is one of the few (GLM is the other) which actually states biological truths, rather then political interpretations.
But professionals aren't asking AI tools about gender politics. They're using them to code and build businesses. I don't care if I'm using a model that has some crazy political takes that I don't agree with as long as it is good at the job it is doing.
Seems if you are okay with it, there's no reason to use anything but the highest effort levels of some other frontier models for the price.
I think Grok provides healthy competition to the other labs, though I do think they bank on groks reputation making it less appealing to many.
https://en.wikipedia.org/wiki/Grok_(chatbot)#Controversies_a...
And here:
https://en.wikipedia.org/wiki/Grok_sexual_deepfake_scandal
I think polarizing is a generous way of describing the problems. My organization has outright banned Grok, because we don't trust SpaceX to hold up to contractual agreements vis-a-vis data-privacy/training. That's the level of reputational damage we're talking about here; and we use Chinese models (*hosted by US providers) for context.
Nobody else wants to be in the blast radius for whatever SpaceX/SpaceXAi does next, or whatever their next controversy is. It is easier, when asked, "Do you use Grok?" just to be able to answer no, instead of having to explain why you aren't embroiled in whatever is going on this week.
Just last week they were fighting Minnesota's law that makes creating this stuff illegal.
The model itself is great though, especially in grok build, which is a really nice harness I find myself preferring these days.
Kind of disappointed by how many people don't see any reason to boycott a model that nudified minors and makes money for a guy that does Nazi salutes.
1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months?
2) Distillation - also implausible for the reason above.
3) Benchmark hacking. AI companies have ways they can dial up performance artificially, and will reach for that to maintain the appearance of parity.
Other reasons?
Edit: Most replies are ignoring timing. It's the near-concurrent release of the same jump in capability that I find suspicious; not the fact that labs can catch up eventually.
But I think we’re discovering that intelligence is about universality, not magnitude. This is analogous to how building a universal Turing machine wasn’t merely a matter of building a calculator that could multiply higher numbers. The difference is that with calculators we consciously theorized about what universal computation would require, then we built one as a step change. Despite it having low memory and slow speeds, the first one built was as theoretically universal as any computer we have today, in terms of the surface of computations it can perform.
With intelligence, it’s turned out to be less discontinuous, which I believe has convinced people that intelligence is a never ending exponential rather than an S curve approaching a horizontal asymptote. I suspect the LLMs we have today are the same kind of thing we will have in 5-10 years, but in 5-10 years we’ll consider them to be fully universal. At that point we’ll still have improvements in tokens per second and volume of context window, but not in capability per token.
The assumed timeline (2 months) is slightly wrong because Fable (Latin) is essentially the same as Mythos (Greek) albeit with protections against cyber and biological misuse.
Mythos (Preview) was publicly announced in April 2026 [1] which means other labs have had 4 months to catch up, not 2 months.
Assuming everyone had access to Mythos from the start, your expression, similar to other folks would have been "Mythos-level intelligence" and not "Fable-level intelligence".
It's not an explanation of why it happens, I am just pointing Fable is not an exception, it has happened with almost every other model release by all these companies over the last 2-3 years.
it used to be snapdragon came out HTC rushed out a janky phone everyone went omg htc is goat, then in the next few weeks and months others would impliment better versions and people would not notice those as much, finally sony would release a polished phone right as the next snapdragon cycle came.
eventually compute gains leveled off and apple won on taste.
nvidia/tpu is the new snapdragon. Anthropic and google both peaked on the first training run on a new tpu cycle.
you should expect amazing things within a few months of each other from everyone with access to chips and willingness to use them on a training run.
We haven't seen willingness from google to do that. So its currently xai,oai,anthropic, and probably soon meta.
I'm not stating this as a fact, but it's a hypothesis I'm keeping in my mix.
It's probably a mix of all of that plus simply always keeping one in the chamber to 1up everyone else when the time is right.
Everyones hyped about the branded phone, but it was the chip that mattered and how fast you rushed a product out after you got it.
Sames true now, except size of training run is also a factor.
Other labs catching up in half a year seems about right.
Last week I gave it a small-sized auth ticket to work on, then stepped away. I came back later that afternoon and found that it had worked for 3+ hours and written 25,000+ lines of code. I skimmed over the code and it looked like a small fix followed by a massive number of additional checks around it, including static analysis tooling.
I gave it to another GPT 5.6 and said "check this code and see if it addresses the ticket". It looked at it and said that 98% of it was garbage and should be thrown away (its own words). I then gave it to Fable, which said it was massively over-engineered. Fable's theory was that the agent implemented the fix first, but then compacted and lost crucial context, forgot what the original task was about, and kept going. After many compaction cycles it was completely lost.
Some people complain that Opus 5 stops before finishing a task. But to me, that behavior is vastly preferable to what GPT 5.6 Sol does.
Explaining it as a difference of effort would explain both.
What are suspicious of? If the timing is similar maybe just everyone already are of similar capabilities and got there at a similar time?
> Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models?
It means Anthropic had no real moat and no real lead. Is that weird to you?
Maybe research is sufficiently public and simple to reproduce or the next steps of how to improve things are sufficiently obvious to the smart people working on frontier AI.
We'll see with 4.6.
But Opus 5/4.8 was better for non-code architecture discussions and general intelligence. However, for the cost, I'd use GPT 5.6 Sol and get much better results. Interestingly, Sol is not great for coding - slow and overengineer stuff if you're not explicit.
My go-to workflow was Sol for planning and Grok for building. But my in my first tests with Grok 4.6, I found it quite good and I'll start using it for both; assuming it's as good at is shows at benchmarks it's unbeatable at cost/time.
I still think that it's very possible Gemini gets its act together and becomes the true competitor to the existing frontier models (on more than just cost). But they sure are taking their time with this one, and recent org changes don't exactly signal confidence
As a designer, I'm always hesitant to believe these statements until there's independent comparisons between the old & new model, as well as comparisons to human made flows. Design can be so subjective that blanket statements like this seem almost useless.
I hope grok4.7 will improve this even more.
My guess is that xai benchmaxxes a lot but fails in actual capacity to produce good models.
- Rocket Design
- Battery Chemistry
- Frontier level AI research
There's no way he's just a guy with a bunch of money paying smart people to do things.
Just imagine how much he's trying to push internally that this new generation of Grok should be spouting his kind of propaganda.
I don't care how smart or cheap the model is if it's run by Musk, I just can't use it.
I've literally never heard someone say they are excited about Musk's CSAM slop bot yet there are like 10 of them here.
By the way if you're looking to get off of Cursor since X.ai bought them, I switched to Zed a few weeks ago and out-of-the-box it does everything Cursor can without charging you their dumbass premium. The transition is pretty much seamless