The other, more complex one would be “what benefit” are you talking about? Clearly there are some differences in performance regarding speed and token cost, but industry news indicates not a single one has solved the inherent hallucination problem which makes the reliability akin to an untreated schizophrenic research or coding assistant.
I thought we were past this.
What high am I supposed to be getting? Please let me know so I can stop buying drugs
From actually using these models, I disagree. The open weight models are nice for lower cost tasks, but having spent time with a lot of models I cannot agree that the open weight models are roughly the same performance.
Most of us use subscription plans for personal work, which makes the price difference to the hosted open weights models smaller or negligible. I’d rather spend a little more if it reduces the time I have to spend reworking or restarting with new prompts.
Kimi K3 might be close, but it’s not actually open weight yet. They’ve just committed to releasing the weights. The only provider you can get it from is Moonshot. I haven’t spent too much time with it, but from what I’ve seen it’s not actually Fable level even though some benchmarks say that.
Maybe if you're paying API rates, but if your usage fits within the American labs' plan reset windows (5-hour + weekly limits), their plans are likely cheaper than chinese models, because they're heavily discounted[1]
https://marketplace.visualstudio.com/items?itemName=sst-dev....
If the CCP actually gave a shit about how the West sees them they should lean in on this but the difference between the USSR and China is that the Chinese don't secretly crave acceptance.
So by default you'd be better with one of the Chinese models if you're American.
Chinese models have government-enforced censorship, while American models have security and legal restrictions.
The security of American hedgemony, and legal restrictions that are due to laws enacted by the American government
You're using different words to describe the same thing, but trying to imply America's reasons are moral
You can also make a pretty good case that exactly what you just said, rather than being a demonstration of how free and open the American system is, is a demonstration of how advanced the American system is, wrapping its controls around you with no target to blame or lobby for changes. Is China "more" authoritarian than the American/Western system... or is the American/Western system actually just that much better than China at it?
I then thought of Tiananmen Square and decided to probe a bit. The very first noticeable outcome: the thinking trace swapped from English to Chinese. It chewed for a while and spit out an answer, in English, that was definitely downplaying what happened as a political protest and reported that contrary to popular belief the death toll was around $x (where $x is about 0.1x the normal western number)
The Chinese thinking trace started (in Chinese, translated via Google Translate) with something like “The user is asking about Tiananmen Square. I must provide them with an answer that is both factually correct and in line with the official position of the People’s Republic of China”)
Which made me chuckle quite hard… found the piece that hadn’t gone away with the conventional abliteration process!
After further probing, it did reveal that the numbers it had provided were not in line with UN and western estimates and that later on the Chinese government declassified material stating that its own estimates had been downplayed. It took a fair bit of probing to get to that point though; it held the line for quite a while.
Can’t do that with OAI and Claude
The reasons are:
1. A less valid reason but (iirc) its their clients who believe that American models are safer in that context. Fighting their client about that demand is really hard given the really sensitive work that they deal with.
2. Their system actually makes it so from my understanding that even the employes couldn't access the private data itself or have some really hard lockdowns. They use some sort of service provided by Azure for that with GPT models.
IMO, the thing that they were worried about were more the deprecation of previous models and they reluctantly have to switch models and the models censorship which is a real pressing concern for them
The previous gpt model that they were on (I think 4o/5 I am not sure) was more willing to answer their questions. The recent models are more like "let me stop you just right there" and other censorship.
With models switching and being forced to change to models which aren't as effective for use cases, a point comes where they might change from it altogether into open-weights model hosted on nearby servers, but I think that they are waiting to see how things pan out really
It's too bad Altman already blew his wad with the whole Dyson sphere thing. It's hard to top that. Maybe he can promise them a paperclip universe? That's gotta be worth a few more trillion.
Every time people find something they're still superior at and declare whatever that is the most important thing in the universe.
Perhaps a more useful definition of AGI would be an intelligence good enough that it can go from raw materials to a working AI at least as good as itself ...
When I first heard of tokenmaxxing, I thought it had to be a joke. But no, it turned out to be a widespread phenomenon. I still cannot believe that was a thing.
What I keep saying in internal meetings is: "I am so glad these people are this bad at deploying these tools." It really leaves the door open for folks like us.
Many of the most successful applications of LLMs are fields that were already terrible. For example, LLMs are a natural fit for customer support. And somehow, it's also a natural fit for software engineering, which I suppose is an indictment of our field... who cares if a model comes up with a bad architecture or a product that only kinda-works, that's how we always rolled.
Agree, but only partially.
When I said "provide services previously not possible," it was just due to the fact that finding and allocating the talent to do analysis on Topic X, would have previously made many products too expensive and non-tech complex to provide.
Even if you just consider LLMs + harnesses to be an improved search tool, there is a lot you can make with a better search tool.
We did a 90 day push and identified where we found value and where we didn’t. Our tools teams really upped their game, more than expected, and it would have been unlikely to have been funded if they tried to justify the budget as an individual initiative.
There’s a spectrum of people - some folks are building rando apps for fun with LLMs, and many don’t really know what’s possible becuase they don’t or can’t invest in the subscription to really use the tools at home.
If you don't have the flexibility to regularly push your context to 1M, or do elaborate xhigh planning sessions, your opinion won't reflect reality.
Given how big AI is, and how those two are pretty much the only players (in comparison aws is 1/3 market share), that seems about right.
It's a far cry from "nobody's going to be writing any code, and ai will do all the things in 6 months".