Fish Audio has raised a $52 million seed round to build AI voice models for creators and enterprises. The round was led by Coreline Ventures and Capital Today, with participation from several other investors.

The Palo Alto startup says more than 8 million people use either the open-source or hosted versions of its models, and that it now generates $21 million in annual recurring revenue. Its technology includes more than 15,000 natural-language controls for steering generated voices, reflecting demand for synthetic speech that is more expressive and easier to direct.

Fish Audio began as a project by former Nvidia researcher Shijia Liao, who trained and open-sourced a voice-generation model after finding existing synthetic voices insufficiently expressive. The company has released five models in the past year, including four speech-generation models and one speech-to-text model. Its newest S2.1 Pro model is available only through a paid API, showing the familiar open-source-to-commercial path for AI infrastructure startups.