90 pointsby nkko11 hours ago10 comments
  • ladyanita222 hours ago
    This is something I've been fantasizing about for long.

    Let's say we took Rust, a language that makes parallelization easier than others (as it helps you avoid some common footguns). How difficult would it be to have a massively parallel computer system made out of many tiny, simple microcontroller-like chips? Let's say we picked many little Risc-V's. Surely this would be an interesting experiment (though I'm not sure whether it'd make economic sense or not...)

    • sigmoid102 hours ago
      It would certainly not make any economic sense, and I guess that's also why noone is seriously looking into stuff like volunteer/enthusiast clusters of home computers to do inference in the same way that e.g. LHC@home works. The main bottleneck for LLMs is still memory bandwidth. Any memory bus not directly soldered on your GPU is terribly slow. That's why one big GPU with twice the VRAM will always perform significantly better than two GPUs with half the VRAM each. And it's also not like you can just solder more memory onto a chip. At modern speeds, the speed of light is a hard limit. For current GDDR7, signals may only travel like 10mm per cycle.

      If you spread such a system out over dozens or hundreds of tiny chips, you'll be wasting most of its resources and lose hard to anyone who built a single chip setup.

    • ur-whale17 minutes ago
      > How difficult would it be to have a massively parallel computer system made out of many tiny, simple microcontroller-like chips?

      It's scaling the communication that becomes hard.

      In this project they daisy-chain SPI. I don't believe that would scale very far.

  • tdhz776 hours ago
    Soon ai in every lightbulb running Kubernetes
    • oneZergArmy2 hours ago
      Praise the Omnissiah.
      • Tade02 hours ago
        With the proliferation of Abominable Intelligence? Quite the contrary!
    • tombert4 hours ago
      You know, I've always liked Futurama but I always kind of thought it was silly that literally everything has an AI and a personality.

      But, you know, I actually think that there might be a logic to it. Economies of scale might mean that almost-literally every computer you buy in the year 3000 has some kind of AI-assistance chip in there, and sure maybe it will have full AI with a personality spitting out one-liners.

      • abroadwin3 hours ago
        Kind of like how disposable vape pens often have a 24 MHz Cortex-M0+ with 3 kB SRAM and 24 kB flash, which would have seemed ludicrous a while back.
      • KeplerBoy2 hours ago
        Change the year 3000 to the 2030s and it might be just as accurate.
  • librasteve2 hours ago
    haha … this is precisely the kind of project that https://bil-lang.org is aimed at: Go for parallel (ie in this case pipeline processing).

    don’t get too excited until we get the TinyGo backend built though ;-)

  • NDlurker5 hours ago
    I'm curious how this would handle grammar checking on a basic word processor. Or maybe generate worlds for small text based games. I have no idea what the capabilities are of a cluster like this.
  • cameron_b8 hours ago
    It is a bit of a bummer to see that the degree of 'compression' makes it a fancy llm noise-maker. It is still charming.
  • matthewfcarlson5 hours ago
    I’m actually working on a small project that’s exactly this! Less quant so it’s only 150M parameters but this is amazing.
  • sneakan hour ago
    Seriously though, what are the low cost chips that can usefully run LLMs? Is a Mac Mini the lowest we can go? Are there iGPUs on mini-itx that can do it, or are there dedicated AI chips that one could turn into a pi HAT?
    • Risse16 minutes ago
      Depends on what you consider to "usefully run LLMs".

      Earlier this year, I bought a mini pc from Aliexpress, specs are roughly Ryzen H255, 24GB LPDDR5, 1TB SSD. This was around 350€ including VAT, customs, shipping etc. I would personally consider this somewhat of a lowest class of useful LLM box. It can run 8B models well, up to somewhere around 24B. I currently run Gemma 4 26B A4B Q5 on it, with MTP, and it is quite slow, but smaller models would run okay on it.

  • sjakati987 hours ago
    Gemma 4 when?
    • nkozyra4 hours ago
      We're gonna need a bigger ESP.
  • aidiscoverywire2 hours ago
    [flagged]
  • nonasking_4 hours ago
    Thanks for sharing. It's fascinating to see a 0.5B LLM being split across seven ESP32s like this.