Fermion Research's Phonon-2 Beats Whisper at 5.21% WER in Just 164 MB
Fermion Research shipped a 164 MB English speech recognition model that matches a 2.5 GB teacher's accuracy and transcribes an hour of audio in 20 seconds on a MacBook Air.
- Fermion Research released Phonon-2, a 164 MB open English ASR model under CC-BY-4.0.
- Averages 5.21% word error across seven Open ASR Leaderboard sets, within 0.25 points of its 2.5 GB teacher.
- Encoder weights stored at one of five learned levels in about 2.1 bits each.
- Runs at 174x realtime on an M5 MacBook Air and 6,680x on an H100 batch 128.
- Install with
pip install fermion-research; CPU and CUDA Docker images available. - Serves an OpenAI-compatible HTTP endpoint and powers the Detta Mac dictation app.
Phonon-2 puts accurate English ASR in a 164 MB download
Fermion Research has released the Phonon-2 weights, an open-weight English speech recognition model with a 164 MB download. It averages 5.21% word error rate across the Open ASR Leaderboard’s seven English datasets. In Fermion’s published comparison, that is the lowest average WER among open models smaller than 900 MB. Lower WER indicates fewer substitutions, deletions, and insertions.
Phonon-2 is distilled from NVIDIA’s Parakeet TDT 0.6B v3 and retains its tokenizer, punctuation, capitalization, and numeral conventions. The compact download is about one-fifteenth the size of the 2,508 MB teacher, making the model practical for local transcription, edge applications, and high-volume batch processing.
Five levels per weight
Phonon-2’s encoder stores each weight as one of five learned values, averaging about 2.1 bits per weight. Six-bit lookup tables support inference. Quantization-aware distillation incorporates those low-bit constraints during optimization, allowing the student model to adapt to the compressed representation.
Across the seven-set benchmark, Phonon-2 trails its full-precision teacher by 0.25 percentage points of average WER. Fermion’s per-dataset results also show 100.8% of the teacher’s word accuracy on parliamentary speech and a lower WER on meeting audio.
| Model | Download | Average WER |
|---|---|---|
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.