TinyTTS is a proof of concept: It asks whether a VITS-style neural text-to-speech model can run on a microcontroller at all. It is a header-only C++ port of tiny-tts. The model has about 1.6M parameters. It has four stages: text encoder, flow, duration predictor and a HiFi-GAN-style vocoder. There is no inference runtime such as TFLite Micro. Everything is hand-written C++, available as an Arduino library, an ESP-IDF component and a desktop CLI.
Memory is not the problem. Weights, a slimmed CMU dictionary and a neural G2P fallback total 4.2 MB. After begin(), PSRAM usage was only about 743 KB.
Speed is the problem. Synthesizing “Hello world!” (1.49 s of audio) took:
| Platform | Time | vs. real time |
|---|---|---|
| ESP32-S3 (baseline) | ~440 s | ~296x slower |
| ESP32-S3 (INT8, SIMD, tiling) | ~34 s | ~23x slower |
| ESP32-P4 (optimized) | ~28 s | ~19x slower |
| Raspberry Pi 4 (4 threads) | ~1.55 s | ~1x |
| Desktop | ~0.76 s | faster than real time |
The model works properly on the desktop and on microcomputers, but even the fastest ESP32 is about 20 times too slow
Where the time goes
The vocoder needs about 440 million multiply-accumulates per second of audio, roughly 90% of the model’s work. The flow accounts for most of the rest. No microcontroller CPU can do that in real time.
The fix: move the heavy part to an FPGA
The TangNanoVocoder project splits the processing up to run on a:
- Microcontroller (e.g.ESP32-S3): text to phonemes, the text encoder and the duration predictor. This is the light part, with 0.66 MB of weights. It produces the latent z_p.
- Sipeed Tang Nano 20K FPGA: the flow and the vocoder. It plays the audio itself over I2S or PWM.

Only the latent crosses the link. It is 32 channels at 86 frames/s in 16 bit, about 5.5 KB/s, so plain SPI or UART is enough. Audio never goes back to the MCU. An 8-bit latent would cost 22 dB at the output, so it has to be 16 bit.
The FPGA runs the model without a CPU in the inner loop. The model is a program image of op descriptors and weights. The FPGA loads it from its own flash at power-up. A hardware scheduler with SDRAM DMA then executes it: a 16-lane convolution engine, LayerNorm, an attention unit, requantize, residual and tanh. The microcontroller only sends sentences.
Results
- 0.91x real time. The FPGA computed 4.05 s of speech in 3.68 s.
- Bit-exact with the C++ fixed-point model (16-bit activations, 8 to 12-bit weights).
- Close to float TinyTTS. The fixed-point model reaches 39.7 dB spectral SNR against float. The vocoder alone reaches 39 dB waveform SNR.
- Pipelined sentences. The FPGA computes the next sentence while the current one plays.
Takeaway
A small neural TTS fits in a microcontroller’s memory, but its compute doesn’t fit. A cheap FPGA board with a purpose-built pipeline closes the gap and keeps the MCU for the light work.
Repos: TinyTTS ยท TangNanoVocoder
0 Comments