Fish Audio's S2.1 Pro Beats ElevenLabs at 11x Lower Cost After $52M Raise

Fish Audio raises $52M seed, launches S2.1 Pro with 5-second voice cloning, 83-language support, and pricing that undercuts ElevenLabs by up to 6x

·
·
AuthorFish Audio
Read2 min
TopicAudio · Business
  • $52M seed raised led by Coreline Ventures and Capital Today; Fish Audio hit $21M ARR and 8M users in its first year.
  • S2.1 Pro launched with 5-second voice cloning, 83-language support, and word-level emotion/intonation control via 15,000+ natural language tags.
  • Pricing undercuts rivals: $15/1M characters vs. ElevenLabs' $60-$165/1M; company claims 2x faster than Cartesia at 1/6th the cost of ElevenLabs.
  • 67% blind-test preference over leading competitors; 56.3 chars/sec throughput, faster than GPT-Realtime-2 and Gemini 3.1 Flash TTS.
  • Enterprise customers include HeyGen, LiveKit, Retell, and Sanas; HIPAA-compliant on-prem deployments now available.
  • S2.1 Pro free via API through August 31; roadmap includes speech-to-speech models and voice-native LLMs.

Fish Audio announced a $52 million seed round and the public launch of its flagship model, S2.1 Pro, a text-to-speech system built around expressiveness, speed, and aggressive pricing. The company reached $21M ARR and 8 million users in its first year.

From a single GPU to $52M

Fish Audio began in the bedroom of co-founder and Chief Scientist Shijia Liao, a former NVIDIA video researcher and lifelong VTuber and anime fan who grew frustrated with monotonous synthetic voices and started training models on a single gaming GPU. The resulting open-source project, Fish Speech, accumulated more than 31,000 GitHub stars and pulled in creators, game developers, and indie builders before becoming the foundation for a real business.

The seed round was led by Coreline Ventures and Capital Today, with participation from 359 Capital, Play Time, HF0, 645 Ventures, Parable, Carya Venture Partners, and Alphalist Partners.

What S2.1 Pro actually does

Most TTS APIs let you pick a voice and adjust speed. S2.1 Pro goes further, offering word-level emotional control through natural language instructions rather than low-level audio parameters. You can write directives like "say this with urgency" or "slow down on this word," which makes it accessible to non-technical users while still giving developers fine-grained API hooks.

  • 5-second voice cloning: Clone a voice from a five-second audio clip in under 15 seconds.
  • Word-level emotion control: Over 15,000 natural language controls for intonation, pacing, and emotional tone at the word level.

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves