Let's say we took Rust, a language that makes parallelization easier than others (as it helps you avoid some common footguns). How difficult would it be to have a massively parallel computer system made out of many tiny, simple microcontroller-like chips? Let's say we picked many little Risc-V's. Surely this would be an interesting experiment (though I'm not sure whether it'd make economic sense or not...)
If you spread such a system out over dozens or hundreds of tiny chips, you'll be wasting most of its resources and lose hard to anyone who built a single chip setup.
It's scaling the communication that becomes hard.
In this project they daisy-chain SPI. I don't believe that would scale very far.
But, you know, I actually think that there might be a logic to it. Economies of scale might mean that almost-literally every computer you buy in the year 3000 has some kind of AI-assistance chip in there, and sure maybe it will have full AI with a personality spitting out one-liners.
don’t get too excited until we get the TinyGo backend built though ;-)
Earlier this year, I bought a mini pc from Aliexpress, specs are roughly Ryzen H255, 24GB LPDDR5, 1TB SSD. This was around 350€ including VAT, customs, shipping etc. I would personally consider this somewhat of a lowest class of useful LLM box. It can run 8B models well, up to somewhere around 24B. I currently run Gemma 4 26B A4B Q5 on it, with MTP, and it is quite slow, but smaller models would run okay on it.