Fish Audio Squeezes 2M Voices and Voice Cloning Into an iPhone App

Fish Audio brings its full voice studio to iPhone with 2M+ voices, 80+ languages, and open-ended emotion tags. Here is what shipped.

·
·
  • Fish Audio launched a native iOS app mirroring its full desktop voice studio.
  • Includes 2M+ voices, multi-speaker dialogue in 80+ languages, and open-ended emotion tags.
  • Free to download with $14.99/month Plus tier and one-off credit packs; requires iOS 16.4+.
  • Powered by S2.1 Pro, which tops Seed-TTS Eval word error rate benchmarks against Qwen3-TTS and MiniMax.
  • Company raised $52M seed in July, hit $21M ARR and 8M users in year one.
  • Grew out of Fish Speech, an open-source project with 31,000+ GitHub stars.

Fish Audio just squeezed its entire voice studio into an iPhone. The Palo Alto startup, which spent the last year turning an open-source hobby project into a serious ElevenLabs competitor, released a native iOS app that mirrors what its desktop and web platforms already offer: a library of more than two million voices, multi-speaker dialogue generation, single-clip voice cloning, and prompt-driven voice design.

The App Store pitch is deliberately simple. Type text, pick a voice, get expressive speech back. Under the hood, this is the same stack powering a company that raised $52 million in seed funding and hit eight million users in twelve months.

What actually ships in the app

The iOS release, published under Fish Audio's legal entity Hanabi AI Inc., weighs in at 82.8 MB and requires iOS 16.4 or later. Download is free, with a Plus tier at $14.99/month or $131.99/year, plus one-off credit packs at $4.99 and $9.99 for one and 2.5 million characters respectively. The mobile build covers:

  • A catalog of 2M+ community and preset voices, searchable by language and style
  • Multi-speaker dialogue generation across more than 80 languages
  • Open-ended emotion and paralinguistic tags (think [whispers sweetly] or [laughing nervously]) inline in the script
  • Voice cloning from a short recording, or generating a new voice from a text prompt
  • Export and share directly to other creative tools

The tag syntax is the interesting bit. On the web, S2 Pro uses simple bracket markers to embed emotional instructions at any position in the text, and supports 15,000+ unique tags rather than a fixed preset list. That same free-form control is what the app exposes as "open-ended emotion tags" in the marketing copy.

The model behind the button

Fish Audio's current production model is S2.1 Pro. It supports 83 languages with automatic language detection, uses natural language control via bracket syntax not limited to a fixed set, and is the company's recommended production TTS system. An older S1 remains available for existing integrations.

The quality claims are non-trivial. On Seed-TTS Eval, S2 posts the lowest word error rate among evaluated systems, beating Qwen3-TTS, MiniMax Speech-02, and Seed-TTS. The company's own blind listening tests report S2.1 Pro was preferred by 67% of listeners over rival models, though those numbers come from company-published data rather than an independent auditor. Voice cloning uses zero-shot inference: the model was trained once on more than 10 million hours of audio across 83 languages and applies that training to new voices at inference time, with a speaker encoder that captures tone, pitch distribution, formant structure, and rhythmic patterns from a short clip.

From a bedroom GPU to $21M ARR

The company's origin story matters for understanding why a mobile app is showing up now. Fish Audio began in the bedroom of co-founder and chief scientist Shijia Liao, a former NVIDIA video researcher frustrated with monotonous synthetic voices. He trained models on a single gaming GPU, and the resulting open-source project, Fish Speech, became one of the most popular voice projects on GitHub with more than 31,000 stars.

That open-source flywheel turned into revenue fast. The $52 million seed round was led by Coreline Ventures and Capital Today, with participation from 359 Capital, Play Time, HF0, 645 Ventures, Parable, and others. It closed as the company marked its first anniversary, having grown from zero to $21 million in annual recurring revenue and more than 8 million users across creators, developers, and enterprises.

Where this lands in a crowded market

The voice AI space is not short on incumbents. ElevenLabs, WellSaid, Cartesia, Speechify, Async, and Krisp are all competing for creator and enterprise budgets. Shipping a polished mobile app is one of the few remaining moves that meaningfully changes who uses these tools, because it opens the door to people who script and record on their phones rather than at a desk.

It also fits the broader roadmap Fish Audio has telegraphed. The company plans to expand its lineup beyond text-to-speech into a full audio-native stack, including voice-native LLMs and speech-to-speech, while deepening developer tooling and integrations with partners like LiveKit and Retell. A mobile front-end is the consumer surface for that stack.

Whether it is worth your time

For anyone already using Fish Audio through the web app or API, the iOS release is mostly a convenience layer. Same models, same voice library, but you can now draft a multi-speaker podcast script on the train and export the audio before you get home.

For ElevenLabs users who are curious, the free tier plus one-off credit packs make it cheap to A/B test the voice quality on your own scripts. The interesting technical bet to evaluate is the tag system: whether prompt-level directions like emotion, pacing, and non-verbal cues actually give you more control than the slider-and-preset interfaces most competitors ship. That is the axis Fish Audio is trying to win on, and the app puts it in reach of anyone with an iPhone.

Comments

avatar