Local AI on your device seems like a much more likely future to me than datacenters in space. For inference at least, training is another story.
Soon decent speed across two Mac Studios with 512GB of RAM.
To really go fast you’d probably have to do PCB layout and do like 256 or 1024 chips in parallel with a fast SRAM aggregation buffer feeding a GPU or TPU rig.
I wonder if anyone is doing this? You would flash in a model and then just run it. It would need RAM for context but much less of it.
idk how people access (soldout) and even afford 512GB RAM MacStudio's. Isn't it $40k or so?
Also its answering the question of what gonna happen if you wake up tomorrow and datacenters are gone. Or internets are gone.
Some people on our globe live in countries with no internet whatsoever. Of course most of them dont have Macbook with 64GB RAM either, but it's much much easier to get than internet connection or rack of GB200.
SOTA LLMs are efficiently compression of all the knowkedge humanity has built. Having ability to run it at home to extract said knowledge is important no matter the speed.
But also that's a pretty extreme hypothetical. Imagine the polymarket on that.
- bunch of people only ~4 years ago
(note. I am not a believer in AGI)
"useful" is highly contextual. The clock of the long "now" is not useful in the sense you mean, to synchronise your wristwatch. I'm still glad it exists.
Kimi Pen Pal. Bring back lettets and postcards. Do OCR, and use one of those 3D printer-like pen plotters write the model output as a letter.
Challenge would be automating the opening and OCR preparation, and the folding and mailing of the return letter. But given it's done commercially it should be possible.
It reminds me of when Willow Garage chose to name their bot the TurtleBot, because if they named it anything else, people would think it was fast and capable. But when they called it Turtle Bot, people just kind of liked it and were satisfied with what it did.
At the level of Kimi 3, I probably can code only about 1,000 good tokens per day, too. (thankfully coding isn't my job)
Need to justify buying an expensive rig that doesn't do what you expected.
Specifically thinking the people they could do something AI with cpu, and realizing it isn't feasible. Happened at my fortune 20 company. They had to get approvals and ofc it was useless. Plenty people tried to explain, but they were the principle engineer, and out ranked everyone.
"It's not going to work", the topic changed, and we never spoke about it again.
16tk/s... Then 3 tks per minute. Then someone else posted 0.3tk/s.
I don't know if I'd call this "running"
UPD: I know it's not the same at all, just the reversal of units that gets me
It’d be like thinking as slowly as Ents talk to each other in Lord of the Rings.