Fish Audio just closed a major seed funding round to accelerate development of its text-to-speech models. The company reports more than 8 million users across its open-source and hosted model offerings, alongside $21 million in annual recurring revenue.

What Fish Audio actually does

Fish Audio builds text-to-speech technology that converts written words into natural-sounding human voices. The company offers both an open-source version of its models and a hosted API product, which lets developers plug voice synthesis directly into their own applications.

Its flagship model family, the Fish-Speech and S2 series, supports more than 80 languages. The company published technical reports on its S2 and S2 Pro models in March 2026, and its benchmarks show a word error rate of 0.54% in Chinese and 0.99% in English.

Real-time streaming and instant voice cloning are two of the platform’s headline features. Fish Audio also runs a community marketplace where developers and creators can share and monetize voice models they have built on top of the platform.