A 46,000-parameter CNN, compiled to WebAssembly, predicts how long every phoneme takes — before any audio exists. Words get start/end times good enough for read-along highlighting. The model file is 183 KB.
A full piper teacher voice (63 MB, patched to expose its internal durations) synthesizes on-device via onnxruntime-web. Words light up as they are spoken, timed by the model itself. SSML works:
Grapheme-to-phoneme and back, in 130+ languages, with the ByT5 tiny models (19 MB int8) running in this tab. The tag picks the language, e.g. .
Three tiers of the floravox ecosystem, all client-side:
| tier | size | where it runs |
|---|---|---|
| timing students (46k params) | 183 KB | this tab, ESP32, anywhere |
| voice students (1.4M params) | ~1.4 MB int8 | ESP32-S3 real-time |
| patched teachers | 20–63 MB | this tab (WASM), desktop |