Antalia-2 Mini Beats 2.38B Models on Turkish Speech at Just 30MB
A 7.6M-parameter Turkish text-to-speech model runs 50x faster than real time on a laptop CPU, with the fewest character errors among 15 Turkish systems benchmarked.
- PatientDesk AI released Antalia-2 Mini, a 7.62M-parameter Turkish TTS model under Apache-2.0.
- Lowest CER (0.19%) among 15 Turkish systems on Freya-TR-Eval, matching models 300x larger.
- Runs 50x faster than real time on an Apple M5 CPU with 2 threads, no GPU needed.
- Flow matching with shortcut self-consistency lets one model sample in 1, 2, 4 or 8 steps.
- Trained on 1,578 hours of synthetic speech generated by VoxCPM2, no real speakers used.
- Install with
pip install antalia-mini; browser demo available on Hugging Face Spaces.
Antalia-2 Mini brings Turkish text-to-speech to CPUs
Antalia-2 Mini is an open-source Turkish text-to-speech model designed for local, CPU-only inference. Its 7.62 million parameters occupy 30.5 MB in fp32, and the project publishes the weights, text normalizer, and training recipe under Apache 2.0.
Turkish-language developers have often had to choose among hosted APIs, multi-gigabyte multilingual models, and compact systems with higher recognition errors. Antalia-2 Mini narrows its scope to one language and one voice, trading flexibility for a small footprint and low latency.
| Parameters | 7.62 million |
|---|---|
| Weight size | 30.5 MB in fp32 |
| Language | Turkish |
| Speaker support | One fixed voice |
| Default sampling | Eight refinement steps |
| Target runtime | CPU, GPU, or browser |
| License | Apache 2.0 |
A 7.62 million-parameter model near the top
Published results on Freya-TR-Eval place Antalia-2 Mini among the most accurate Turkish systems measured by word error rate and character error rate. WER and CER compare generated speech with the intended text after automatic transcription, with lower percentages indicating fewer recognition errors.
| Evaluation | WER | CER | Placement |
|---|---|---|---|
| Default settings | 0.97% | 0.19% | Tied with EMA Lightning on WER; lowest CER |
| Public benchmark rerun | 1.05% | 0.20% | Second of 15 on WER; first on CER |
The public rerun placed Antalia-2 Mini 0.01 percentage points behind EMA Lightning on WER, a difference the project describes as within seed-to-seed variation. The same leaderboard includes Trendyol-TTS at 2.38 billion parameters, Anka TTS at 336 million, Chatterbox Multilingual at 500 million, and Gemini 3.8 Flash TTS. Each recorded a higher CER than Antalia-2 Mini.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.