Kyutai Retrains PocketTTS With Simpler Loss, Hitting 0.90% Word Error Rate
Kyutai trained its 100M-parameter on-device text-to-speech model with a new one-step generative objective, hitting sub-1% WER without Jacobian-vector products.
- Kyutai retrained PocketTTS with a drifting objective, hitting 0.90% WER on LibriSpeech test-clean.
- First speech model and first autoregressive model trained with the drifting one-step generative loss.
- Key trick: a learned kernel temperature that stabilizes training on streaming speech latents.
- Removes the Jacobian-vector products that the previous LSD training recipe required.
- 24-layer teacher trained 400k steps, then CFG-distilled to a 6-layer student.
- Full write-up and ablations at kyutai.org; based on Deng et al.
Kyutai has published a technical account of retraining PocketTTS, its 100 million-parameter, CPU-only text-to-speech model, with a one-step generative objective called drifting. The resulting checkpoint records 0.90% word error rate on LibriSpeech test-clean and preserves PocketTTS features such as voice cloning and streaming generation. Its main training advantage is simpler optimization: drifting removes the Jacobian-vector products required by Kyutai’s previous loss.
Kyutai describes PocketTTS as the first speech model, and the first autoregressive model of any kind, trained with the drifting objective introduced by Deng et al. Adapting the method to streaming speech depended on making its kernel temperature learnable.
One latent every 80 milliseconds
During generation, PocketTTS emits one continuous audio representation, called a latent, every 80 milliseconds. A causal transformer tracks the preceding text and audio context, while a compact sampler head transforms Gaussian noise into the next latent in one forward pass. This single-step sampler helps the 100 million-parameter model run on a laptop CPU.
The previous sampler used LSD, a one-step flow-matching loss that learns how to transport noise toward speech data. Training it requires Jacobian-vector products, extra derivative calculations that add implementation complexity and computational cost. Drifting reaches comparable benchmark quality without those calculations. Kyutai has not published wall-clock or memory comparisons, so the demonstrated gain is a simpler training procedure rather than a quantified training-speed increase.
Attraction, repulsion, and temperature
During training, the drifting sampler draws 64 candidate particles for each latent position. A kernel-derived drift field pulls every particle toward the target speech latent and pushes it away from the other particles. The attraction teaches the correct output, while the repulsion preserves variation and limits mode collapse, where the sampler produces nearly identical outputs for different noise inputs. Inference still generates each latent with one sampler pass.
Kernel temperature controls the scale and strength of interactions among those particles. Kyutai updates the temperature when the configured temperature-loss weight is positive, then preserves the learned value when that weight is zero during later fine-tuning. Fixed, manually selected temperatures proved fragile for autoregressive speech; learning the value made training stable enough to complete.
Quality survives the smaller model
Kyutai evaluates intelligibility with word error rate, or WER, by transcribing generated speech and comparing it with the requested text. Lower values are better. UTMOS is an automated estimate of perceived audio quality, with higher values indicating better predicted quality.
| Checkpoint | Training loss | WER | UTMOS |
|---|---|---|---|
english_2026-09 |
LSD | 0.90% | 4.36 |
english_drifting_26-09 |
Drifting | 0.90% | 4.37 |
Training proceeds in two stages. Kyutai first trains a 24-layer teacher for 400,000 steps at a constant learning rate, reaching 1.01% WER and 4.26 UTMOS on LibriSpeech test-clean. Depth distillation then compresses the teacher to six layers. The released student also receives classifier-free guidance distillation from the drifting teacher at guidance scale 2.0, folding the teacher’s guided behavior into the smaller network so runtime generation remains single-pass.
The published table covers clean English audiobook speech and uses an automated quality predictor. It contains no speaker-similarity score for voice cloning or results for noisy text, accents, specialized vocabulary, and long-form stability. Production evaluations should therefore include representative voices and content alongside the reported benchmark.
Autoregression raises the stakes
Autoregressive speech provides a demanding test for a one-step sampler because every generated latent becomes context for the next one. Small errors can accumulate across an utterance and degrade pronunciation, timing, or stability. The reported 0.90% WER provides evidence that drifting can remain stable across a generated sequence, extending earlier demonstrations centered on one-step image generation.
A separate theoretical analysis connects drifting to established score-matching methods. Under a Gaussian kernel, the drift operator equals a score difference between smoothed distributions. A score is the gradient that points toward regions of higher probability, the same mathematical object used in diffusion training. This relationship gives researchers familiar diffusion concepts for reasoning about drifting’s behavior.
Run the drifting checkpoint
The new weights ship through the existing PocketTTS package. Selecting english_drifting_26-09 changes the checkpoint while retaining the same model interface and inference path.
- 100 million parameters under the MIT license
- CPU inference using two cores
- About 200 milliseconds to the first audio chunk
- Approximately six times faster than real time on a MacBook Air M4
- Voice cloning from a short WAV sample
- Streaming output and support for arbitrarily long inputs
- Community ports for WebAssembly, ONNX, and MLX
Two commands to start
pip install pocket-tts
pocket-tts generate --language english_drifting_26-09 --text "Hello world"Existing deployments can test the checkpoint by changing their model selection and retaining the current inference code. For teams training one-step generative heads, Kyutai’s recipe offers the broader result: replace LSD and its Jacobian-vector products with drifting, learn the kernel temperature, and preserve single-pass generation after distillation.