I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture.
In the recent leaked DeepSeek investor meeting, they also mentioned only having a 20K GPU cluster (unclear if NVIDIA, or Huawei).
> A person familiar with Moonshot’s procurement strategy confirmed that the company does indeed have a channel for accessing Blackwell processors via Southeast Asia. They didn’t specify whether this was a rental channel — which is legal, in most cases — or direct purchases, which are a breach of US regulations. The Information reported this week that Moonshot is seeking additional Blackwell processors to train its next model.
> The 20,000 chips Moonshot accesses via Alibaba, meanwhile, are from Nvidia’s earlier generation of Hopper products, the people familiar with the agreement said.
So the 20k GPUs from Alibaba is only a lower bound on how many you need to train a model like Kimi K3.
The advantage of having more GPUs in any case is not so much that you can train bigger models, but that the turnaround time is faster, so you can run more experiments to dial in training choices. It's entirely possible that Musk has more than enough compute, but can't hire the talent to run all those experiments. (That would also explain why he has excess capacity he can rent to Google.)
A few quotes from the transcript:
> Our current computing capacity is approximately 20,000 H-equivalent units, most of which have just arrived within the past month or two
> Regarding the Huawei 950, Huawei currently provides us with 16,000 SIM cards
> A Huawei 950 [cluster] with 16,000 cards is equivalent to only a B-series card [cluster] with 4,000 cards.
(We've changed the title to what the article says now.)
When OAI released gpt-oss it was released as an mxfp4 checkpoint.
OAI, Ant, et al are also obviously employing QAT.
In this sense, any advance in intelligence is a performance improvement and vice versa
Moreover, it’s very plausible (and expected) to use multiple clusters and GPU types for RL rollouts which could very well not be included in this count.
No part of this pipeline is fixed in stone.
I think the word "distillation" needs to be used a bit more selectively here. If their pre-training run was complete before Fable was released that implies that ZERO Fable data went into the base model. Perhaps the timeline allows for a few weeks at best of incremental post-training on some limited amount of Fable data, but calling this "distillation" seems a bit dramatic especially given the redacted outputs that would have been available. A more factual speculation would just be that they may have had time to post-train using a limited amount of Fable output in some fashion (LLM as judge? SFT? Who knows ...).
Are you sure they are using all of their compute on training? Didn't they rent out a ton to other AI companies?
the only way is down for the massive valuations and 'a.i' revenue projections.
Can’t possibly be intentional media strategy by a geopolitical target, right? If its going to negatively impact valuations and revenue of major rivals, seems like a desirable strategy?
We just saw OpenAI take the steps to significantly lower the cost of one of their models, which confirms that at least one western lab has a large margin on inference, not a large inference cost.
Meanwhile, we don’t know how many experiments these east/west labs are performing relative to each other. We also know that many western labs have a whole portfolio of models too, which is product breadth not necessarily waste.
Its all a fucking capitalistic farce to display to other rich elite that "Look at how much clout I have! I can make these peons dance around and do my bidding! Im a slave-owner!"
https://infosec.exchange/@david_chisnall/116991627711001827
He noted that that Silicon Valley doesnt really want to SOLVE problems. They want to find already-solved problems with problem matching. And of course, we just throw more people and more compute instead of optimization and understanding.
The Chinese are being actively constrained with bullshit politics around a second Red Scare moment. And, well, they're winning. A lot.
I just bought 2 switches, 48 port 10GbE with 4x QSFP+ at 40Gb fiber. $110 each.
You can even get 24 port QSFP+ @100Gb networking devices for $350.
Yeah while ram and gfx is $$$$$, networking is rock bottom prices..
This feels like a very American way of designing things - just throw more horse power at it, bigger is better! The rest of the world is usually a bit more resource constrained and efficient at using those resources.
See also Mustangs vs German sports cars, giant American fridges, giant American suburban McMansions vs livable cities etc etc.
Source?
The goal will be to develop smaller models with more efficient architectures, that have similar or even better performance than larger models.
Continual learning tends to imply individualized models, else there is no data privacy (the secrets learned on the job at your company now being available to your competition), which really turns the current AI business model of a single centralized model served to everyone on it's head. If every customer has a different model that essentially means the end of batch processing with the same weights loaded into the GPU.
The direction this suggests is a move away from centrally served common models to locally served individual ones, which generally requires them to be smaller, even if some larger companies may be willing to invest in beefier hardware.
I think this is at least in part why the AI companies are trying NOT to implement true continual learning and see if they can instead finesse it by implementing continual compacted(?) memorization instead, since then it's "just" additional context that needs to be recalled and fed into every request, not weights that need feeding into the GPU. I don't think memorization is any substitute for learning, especially learning of practiced skills, but since it's far easier to implement, and non-disruptive to the cloud-based API business model, this is what we will see first.
The recent news of NVIDIA' investment in Sutskever's SSI has a tiny hint of this also, talking about SSI advising on NVIDIA's future architectural direction - apparently pushing it in a different direction than current (cloud-based, pre-trained) models. NVIDIA may be quite happy to see a move towards local models.
The frontier labs would be well served in carving out 20k sub-clusters and giving research teams carte blanche in building things with radically different architectures - with full permission to distill whatever they want from the flagship models. We'd expect to see more product lines that feel "different" from the flagship models if this were already being done.
If they got a lot cheaper and only a little dumber, and we got a little smarter about how we use them, they could appear to the bystander as much more useful than they are.
I am adding multi-modality to https://github.com/guilt/TinyToT, and I see that dis-aggregating capabilities, very similar to how our own sensory organs work, seems to be paying off quite well.
A VAX 11/780 was good, but an 80386 was a lot better, since the latter could run on 3 AA batteries and the former needed 6,000 watts of 3 phase.
isn't that, in the end, the case with all/most technologies?
i can get a petaflop of compute capacity in a dgx spark for relatively cheap nowadays. that used to be a whole supercomputer like 20 years ago.
Someone will find a way to make cheaper compute, and since nothing fundamental changed there (LLM didn't change how silicon were made), that's bound to happen.
Either they have Blackwell with native 4-bit floating math, or they use have Chinese domestic NPU that support mxfp4 natively.
The article’s statement does not make sense.
Personally, I think 20K nvidias is a stop gap solution because they really don't have the capacity to serve their models to earn any money right now.
From the article:
A person familiar with Moonshot’s procurement strategy confirmed that the company does indeed have a channel for accessing Blackwell processors via Southeast Asia. They didn’t specify whether this was a rental channel...
Probably either the author or the people they interviewed got the specific GPU details mixed up or wrong. Maybe it was B200s.
(Kimi K3 tech report section 4.1.1 https://arxiv.org/pdf/2607.24653)
If Alibaba can't bring the chips into China, but can buy a whole bunch in Thailand or Singapore or whereever (or secure exclusive rights via JV partners where relevant) and then just provide them as a service to its customers in China - what is the point? I'm sure many customers would actually prefer such an arrangement.
[1] https://www.tomshardware.com/tech-industry/artificial-intell...
I wonder how much spending is motivated by the phenomenon of sudden emergent performance in LLMs. Clearly some people who are smarter than me expect something like emergent AGI, or at least they think the odds justify spending whatever it takes to see if that would happen.
That leaves a lot of room for efficient aggressive followers.