7-node ESP32-S3 cluster running a 0.4B LLM via 1.58-bit (BitNet) ternary quantization over SPI daisy-chain - Low-Zi-Hong/ESP32s3-LLM-Cluster

150 points•nkko•2 days ago•31 comments•

31 comments

ladyanita222 days ago
This is something I've been fantasizing about for long.

Let's say we took Rust, a language that makes parallelization easier than others (as it helps you avoid some common footguns). How difficult would it be to have a massively parallel computer system made out of many tiny, simple microcontroller-like chips? Let's say we picked many little Risc-V's. Surely this would be an interesting experiment (though I'm not sure whether it'd make economic sense or not...)

sigmoid102 days ago
It would certainly not make any economic sense, and I guess that's also why noone is seriously looking into stuff like volunteer/enthusiast clusters of home computers to do inference in the same way that e.g. LHC@home works. The main bottleneck for LLMs is still memory bandwidth. Any memory bus not directly soldered on your GPU is terribly slow. That's why one big GPU with twice the VRAM will always perform significantly better than two GPUs with half the VRAM each. And it's also not like you can just solder more memory onto a chip. At modern speeds, the speed of light is a hard limit. For current GDDR7, signals may only travel like 10mm per cycle.

If you spread such a system out over dozens or hundreds of tiny chips, you'll be wasting most of its resources and lose hard to anyone who built a single chip setup.

Mojo might be what you want, especially with "MAX" which is their AI modeling framework, where in other languages you need NVidia's libraries, or AMDs, etc the Max libraries just let you talk directly to the GPU / CPU / ASIC with Mojo. I think Mojo is very underrated in this space right now, but assuming they don't mess it up, it could be a major contender in AI. In theory, if someone releases a board, and Modular (company that 'owns' Mojo) adds it to Max, you're basically in the green to experiment as much as you want.

https://max.modular.com/

tyingq1 day ago
Not exactly what you're describing, and not shipping yet, but a cluster in a box. 8 cores/node, 8 nodes.

https://milkv.io/cluster-08

akavel2 days ago
See GreenArrays' 144-core Forth chips by Chuck Moore.
tdhz772 days ago
Soon ai in every lightbulb running Kubernetes
oneZergArmy2 days ago
Praise the Omnissiah.
Tade02 days ago
With the proliferation of Abominable Intelligence? Quite the contrary!
tombert2 days ago
You know, I've always liked Futurama but I always kind of thought it was silly that literally everything has an AI and a personality.

But, you know, I actually think that there might be a logic to it. Economies of scale might mean that almost-literally every computer you buy in the year 3000 has some kind of AI-assistance chip in there, and sure maybe it will have full AI with a personality spitting out one-liners.

abroadwin2 days ago
Kind of like how disposable vape pens often have a 24 MHz Cortex-M0+ with 3 kB SRAM and 24 kB flash, which would have seemed ludicrous a while back.
KeplerBoy2 days ago
Change the year 3000 to the 2030s and it might be just as accurate.
_joel1 day ago
How many rollingupdate pods does it take to change a lightbulb?
jagged-chisel1 day ago
Eventually.
cameron_b2 days ago
It is a bit of a bummer to see that the degree of 'compression' makes it a fancy llm noise-maker. It is still charming.
librasteve2 days ago
haha … this is precisely the kind of project that https://bil-lang.org is aimed at: Go for parallel (ie in this case pipeline processing).

don’t get too excited until we get the TinyGo backend built though ;-)

NDlurker2 days ago
I'm curious how this would handle grammar checking on a basic word processor. Or maybe generate worlds for small text based games. I have no idea what the capabilities are of a cluster like this.

Read the full thread on Hacker News →

Related stories