Fish Audio's S2.1 Pro Beats ElevenLabs at 11x Lower Cost After $52M Raise

Fish Audio raises $52M seed, launches S2.1 Pro with 5-second voice cloning, 83-language support, and pricing that undercuts ElevenLabs by up to 6x

·
·
  • $52M seed raised led by Coreline Ventures and Capital Today; Fish Audio hit $21M ARR and 8M users in its first year.
  • S2.1 Pro launched with 5-second voice cloning, 83-language support, and word-level emotion/intonation control via 15,000+ natural language tags.
  • Pricing undercuts rivals: $15/1M characters vs. ElevenLabs' $60-$165/1M; company claims 2x faster than Cartesia at 1/6th the cost of ElevenLabs.
  • 67% blind-test preference over leading competitors; 56.3 chars/sec throughput, faster than GPT-Realtime-2 and Gemini 3.1 Flash TTS.
  • Enterprise customers include HeyGen, LiveKit, Retell, and Sanas; HIPAA-compliant on-prem deployments now available.
  • S2.1 Pro free via API through August 31; roadmap includes speech-to-speech models and voice-native LLMs.

Fish Audio announced a $52 million seed round and the public launch of its flagship model, S2.1 Pro, a text-to-speech system built around expressiveness, speed, and aggressive pricing. The company reached $21M ARR and 8 million users in its first year.

From a single GPU to $52M

Fish Audio began in the bedroom of co-founder and Chief Scientist Shijia Liao, a former NVIDIA video researcher and lifelong VTuber and anime fan who grew frustrated with monotonous synthetic voices and started training models on a single gaming GPU. The resulting open-source project, Fish Speech, accumulated more than 31,000 GitHub stars and pulled in creators, game developers, and indie builders before becoming the foundation for a real business.

The seed round was led by Coreline Ventures and Capital Today, with participation from 359 Capital, Play Time, HF0, 645 Ventures, Parable, Carya Venture Partners, and Alphalist Partners.

What S2.1 Pro actually does

Most TTS APIs let you pick a voice and adjust speed. S2.1 Pro goes further, offering word-level emotional control through natural language instructions rather than low-level audio parameters. You can write directives like "say this with urgency" or "slow down on this word," which makes it accessible to non-technical users while still giving developers fine-grained API hooks.

  • 5-second voice cloning: Clone a voice from a five-second audio clip in under 15 seconds.
  • Word-level emotion control: Over 15,000 natural language controls for intonation, pacing, and emotional tone at the word level.
  • 83 languages: Multilingual speech generation with improved quality, lower latency, and higher throughput than its predecessor.
  • Speed: 56.3 characters per second, ahead of GPT-Realtime-2 (45.8 chars/s) and Gemini 3.1 Flash TTS (25.3 chars/s).
  • Listener preference: Preferred by 67% of listeners over leading competitors in blind tests.

The pricing play

ElevenLabs' free tier gives you roughly 6 to 10 minutes per month before the paywall kicks in, OpenAI TTS has no free tier, and Google's Gemini TTS charges from the first token. Fish Audio is betting that being the cheapest credible option at scale is a durable advantage.

S2.1 Pro costs $15 per million characters. ElevenLabs charges $60 to $165 per million, making Fish Audio 4 to 11 times cheaper depending on the tier. To lower the barrier further, Fish Audio is making S2.1 Pro free via API through August 31, 2026, so developers can build and evaluate without a usage countdown.

Who is already running it

HeyGen, Sanas, and Plaud are among the organizations already using Fish Audio's enterprise APIs. The use cases span very different requirements. "Every enterprise has different use cases and different preferences. For example, companies like HeyGen, which use our voices to power AI avatars, want realism in voices; a gaming studio would want expressive voices for their characters; and voice agent companies like LiveKit want more natural-sounding and low-latency voices," said CEO and co-founder Rissa Cao.

For regulated industries, Fish Audio offers on-premises deployment, zero-data-retention policies, and HIPAA-compliant configurations. The company previously ran lean on creator and developer subscriptions; investor interest and demand from regulated industries pushed it toward raising capital and building a sales team.

The competitive landscape

The TTS market is crowded and well-funded. ElevenLabs raised $500M at an $11 billion valuation earlier this year. Cartesia has carved out a niche on raw latency for real-time voice agents, where a 90ms response time creates a meaningfully different user experience than 300ms. Fish Audio sits between them: stronger expressiveness and cloning fidelity than most, at a fraction of ElevenLabs' price, though Cartesia's latency lead in real-time agent scenarios remains a real gap.

Fish Audio has shipped five models in the past year: four speech-generation models and one speech-to-text model. Three of the speech-generation models are open-sourced; S2.1 Pro is available only through the paid API. The open-source flywheel is deliberate: attract developers with free weights, convert them to the hosted API when they need reliability and scale.

The voice consent problem

Fish Audio's community voice library hosts over 2 million voices and has drawn controversy. Some creators alleged their voices were uploaded without consent. The company has since automated its takedown process: submit a short voice sample or a contract proving ownership, and the voice is removed in under three minutes. The underlying tension persists, though. Until a creator discovers their voice is on the platform, it remains available.

Lead investor Osuke Honda of Coreline Ventures addressed the issue directly, calling for "consent, transparency, and attribution" to be built into the product rather than treated as afterthoughts, and advocating for verified voice ownership and revenue-sharing models. The legal risk extends across the sector. In 2026, seven Pulitzer- and Emmy-winning journalists sued ElevenLabs for allegedly using their voices without consent to train voice models, a case that puts training data provenance under scrutiny industry-wide.

What the $52M funds

Fish Audio has outlined four priorities for the capital:

  1. Expand beyond TTS into a full audio-native stack, including voice-native LLMs and speech-to-speech models
  2. Build out an enterprise sales team to capture more regulated-industry customers
  3. Deepen developer tooling and integrations with partners like LiveKit and Retell
  4. Release an audio understanding model later this year

The speech-to-speech model is the most consequential item on that list. It would put Fish Audio in direct competition with real-time voice agent infrastructure, where Cartesia currently leads on latency. If Fish Audio can close that gap while holding its expressiveness and pricing advantages, the competitive picture shifts considerably. For now, S2.1 Pro is free via API through August 31, giving any team building a voice-enabled product a low-friction way to compare it against whatever they are currently paying for.

Comments

avatar