Fish Audio announced a $52 million seed round on July 28, 2026, led by Coreline Ventures and Capital Today, exactly one year after the company launched. The Palo Alto startup reports $21 million in annual recurring revenue and more than 8 million users, and shipped a new flagship model, S2.1 Pro, alongside the round. The most telling detail sits in the release strategy: after open-sourcing three of its four speech-generation models, Fish Audio is keeping S2.1 Pro closed, available only through its paid API.

That last part is the story I want to write about, because I called Fish's S2 Pro the default self-hosting recommendation in my open TTS roundup last week, and the successor to that model will never appear on a list like it.

How did Fish Audio grow this fast?

The origin story is genuinely good, which is presumably why the PR leans on it so hard. Shijia Liao, a former NVIDIA video researcher and self-described VTuber and anime fan, got annoyed enough at flat synthetic voices to train his own model on a single gaming GPU, then open-sourced it. That project, Fish Speech, now has more than 31,000 GitHub stars and became the on-ramp for indie developers, game designers, and creators who eventually turned into paying customers of the hosted platform.