- Sources: Moonshine Micro (GitHub), HN discussion
- Summary: Moonshine AI published Moonshine Micro, an embedded voice toolkit that runs voice activity detection, command speech-to-text (a SpellingCNN model), and neural text-to-speech entirely on a microcontroller, using the roughly 0.80 USD Raspberry Pi RP2350 as reference hardware. The project reports a full-stack footprint of about 3.6 MB flash and 468 KB SRAM (VAD about 89 KB flash, STT about 1.3 MB, TTS voice pack about 1.8 MB) and a classify-plus-speak latency of about 0.7 to 1.0 seconds on the RP2350. The core code and the included SpellingCNN and TinyVadCNN models are MIT licensed. It lands the same day as Transcribe.cpp (see Developer tools), a related local-inference release.
- Comments: HN commenters shared a demo video and an OpenAI and ElevenLabs compatible HTTP wrapper built around the models.
- Why it matters: A complete offline speech interface at single-digit-megabyte footprint moves usable ASR and TTS onto commodity microcontrollers with no network and no host CPU.
send feedback on this story