Building local conversational voice agents usually comes with a frustrating trade-off: multi-gigabyte models sound great but add 1–2 seconds of latency, while tiny micro-models often suffer from terrible Word Error Rates (skipping words or mumbling consonants).
To see how much quality could fit into a micro footprint, I built Vaniq-Edge—an end-to-end, ~8.5M parameter standalone local TTS engine.
Text goes in, 24kHz mono audio comes out. No external vocoders or second-stage models running in the background.
Total Parameters: ~8.5M
Model Size: 34.3 MB (FP32)







