Resume
-
This 7-node ESP32-S3 cluster runs a ~0.4B parameter LLM via SPI chain
-
The master node runs the BPE tokenizer and the INT4 implementation; 6 compute nodes process transformer layers
-
Slow but fun: about 9 seconds per mark — not for heavy use, but great as a DIY project
We got to the point with LLM where we could squeeze them into the ESP32. Of course, it won’t take away the bigger, more powerful LLM models running on powerful hardware any time soon, but they can be a lot of fun as something fun to do at home to handle simple queries. You can even add an external LLM to one; our own Adam Conway connected his local LLM to a $30 ESP32 display and it generates a new screen for every question he asks.
But what if you connect multiple ESP32s to a cluster and run LLM from it? Someone made a ~0.4B LLM setting that does exactly that. While it won’t win any awards for speed, it’s a great DIY project you can do at home.
This seven-node ESP32 cluster runs the LLM in tandem
It’s not very fast, but it’s very cool
in a post on the site ESP32 subreddituser Major-Nebula1743 shared details of his ESP32 AI cluster. He compiled the previous draft he received 56M parameter model powered by ESP32 via ESP-NOWand with the power of integrating more ESP32s, they were able to create something bigger.
How does this work:
– 1 master node + 6 compute nodes (all ESP32-S3 N16R8)
– I use SPI network to connect ESP32 devices for high speed communication. (No WiFi overhead)
– The master node handles the BPE tokenizer and INT4 token embeddings.
– Calculate each process, 4 layers of Transformer blocks nodes and move the intermediate X to the next node to calculate further layers.
So will this replace ChatGPT? Not at all. It is very slow; It takes about 9 seconds to go through a single token in the 0.4B parameter model, so you won’t be using this to vibcode anything anytime soon. However, as a great little project, it definitely gets the job done. You can read more about this project and even build your own the project’s GitHub repo.
