floravox — every model on this page runs in your browser tab

A 46,000-parameter CNN, compiled to WebAssembly, predicts how long every phoneme takes — before any audio exists. Words get start/end times good enough for read-along highlighting. The model file is 183 KB.

A full piper teacher voice (63 MB, patched to expose its internal durations) synthesizes on-device via onnxruntime-web. Words light up as they are spoken, timed by the model itself. SSML works:

<break time="300ms"/> · <prosody rate="slow"> · <mark name="x"/>

first click loads the voice, then tap play

Grapheme-to-phoneme and back, in 130+ languages, with the ByT5 tiny models (19 MB int8) running in this tab. The tag picks the language, e.g. eng-US, spa-ES, deu.





Three tiers of the floravox ecosystem, all client-side:

tiersizewhere it runs
timing students (46k params)183 KBthis tab, ESP32, anywhere
voice students (1.4M params)~1.4 MB int8ESP32-S3 real-time
patched teachers20–63 MBthis tab (WASM), desktop

Engine: Rust→WASM for the student, onnxruntime-web for the teacher. Source: github.com/AACTools/floravox