Voice emotion control is leaving the markup. Kakao's Kanana-o model, detailed on its tech blog on August 4, 2026, now takes plain-language instructions like "read it in a sad voice" or "read it in a Gyeongsang dialect" and reflects them in speed, volume, pitch, emotion, intonation, and intensity. On the Korean InstructTTSEval benchmark it scores 94.50, ahead of OpenAI's GPT-4o-mini-tts at 91.10 and just behind Google's Gemini 2.5 Flash Preview TTS at 95.38.

I work at Speechify, on the SpeechifyAI API side, so a post arguing that the markup layer is thinning is mildly inconvenient for my employer's docs team. Read it with that in mind. I think the direction is real anyway, and the trade-offs are worth thinking about before you build on it.

What can Kanana-o actually do?

Kanana-o is Kakao's in-house omni model, developed by its Unified Foundation Model team. The August 4 update is about the speech layer: the model reads text naturally, and it now also follows delivery instructions written as ordinary sentences. "Read it very quickly." "Read it in a low voice." "Read it in a Gyeongsang dialect." Kakao says the generated speech reflects speed, volume, and pitch, plus emotion, intonation, and intensity.