EMA Lightning Beats ElevenLabs on Turkish Speech in Just 34 MB
A 8.6M parameter Turkish text-to-speech model beats ElevenLabs v4 and a 2.38B competitor on accuracy while running 440x faster than real time.
- EMA Lightning is an 8.6M parameter Turkish TTS model, Apache 2.0, about 34 MB total.
- Hits 0.92% WER on Freya-TR-Eval, beating ElevenLabs v4, Gemini 3.8, and 2.38B Trendyol-TTS.
- 440x faster than real time single request, 1,316x batched on an RTX 4090.
- First audio in 3.86 ms; about $0.0085 per million characters of generated speech.
- DiT with flow matching, distilled to 4 steps via DMD2, plus shared modulation and a windowed Gaussian aligner.
- Install with
pip install ema-lightning; source on GitHub, Turkish-only, single voice.
EMA Lightning Packs Turkish TTS Into 34 MB
Canberk Aslan has released EMA Lightning, an 8.6 million-parameter text-to-speech model for Turkish. According to benchmarks published with the model, it records a lower word error rate than ElevenLabs v4, Gemini 3.8 Flash, and the 2.38 billion-parameter Trendyol-TTS while running locally on a CPU or GPU.
The Apache 2.0 release includes model weights, source code, and a Python package. Its 34 MB stack generates one fixed Turkish voice and requires no cloud API, making it relevant to developers who need low-cost, offline speech generation with predictable latency.
Benchmark Gains With Clear Boundaries
On Freya-TR-Eval, a benchmark containing 495 Turkish sentences, the model card reports a 0.92% word error rate. Trendyol-TTS is roughly 277 times larger by parameter count and records 0.95% in the same comparison.
| System | Parameters | WER | UTMOS |
|---|---|---|---|
| EMA Lightning | 8.6M | 0.92% | 3.30 |
| Trendyol-TTS | 2.38B | 0.95% | 3.83 |
| Gemini 3.8 Flash | Undisclosed | 1.35% | 3.52–3.61 |
| ElevenLabs v4 | Undisclosed | 1.43% | About 3.30 |
Word error rate measures how accurately a speech recognizer can recover the generated text, with lower values indicating fewer transcription errors. It captures intelligibility rather than voice preference or emotional range. The results come from the author’s 495-sentence comparison, so broader and independently reproduced evaluations would provide stronger evidence across domains and sentence styles.
UTMOS is an automated estimate of perceived speech quality, with higher scores indicating greater predicted naturalness. EMA Lightning’s 3.30 score matches the reported ElevenLabs result and trails Gemini and Trendyol-TTS. No human listening study is included in the published comparison.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.