The market for AI-generated voice models is massive. Creative use cases require AI voice models to be more expressive, while enterprises looking to automate customer support and sales ops need them to be more steerable.
Palo Alto-based Fish Audio wants to cater to all of those use cases with its library of more than 15,000 natural language controls. Since launching last year, the startup today has more than 8 million people using the open-source or hosted versions of its models, and now generates annual recurring revenue of $21 million.
To continue building on that traction, the startup on Tuesday said it has raised $50 million in a seed round that was led by Coreline Ventures and Capital Today. The funding also saw participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0.
Fish Audio started as a small project by former NVIDIA researcher Shijia Liao, who, frustrated by non-expressive synthetic voices available on the market, trained a voice generation model on a single GPU, which he open-sourced. The Fish Speech repository on GitHub now has more than 31,000 stars, and is used by indie developers, video game designers, and creators.










