• Sources: project repository, HN 49050512
  • Summary: A project published on GitHub runs a 28.9M-parameter model trained on the TinyStories dataset on an ESP32-S3 with 512KB SRAM, 8MB PSRAM, and 16MB flash. The repository reports 4-bit quantization producing a 14.9MB model and roughly 9.5 tokens per second end to end. It uses per-layer embeddings, the technique from Google's Gemma models, so that about 25M parameters of embedding table stay in slow flash while computation runs from fast memory. The author states the model generates short stories and does not answer questions, follow instructions, or write code, and describes the parameter count as roughly 100 times the previous comparable microcontroller result. The measurements are the author's own and no license is stated in the repository content this run could read.
  • Why it matters: Moving the embedding table into flash and keeping only the compute path in RAM is a transferable trick for anyone fitting a model onto a device where flash is plentiful and RAM is not.

send feedback on this story