Kokoro-82M Clones Any Voice From a 3-Second Clip for Under $20
A tiny 24MB adapter turns Kokoro-82M into a zero-shot voice cloner that enrolls a reference clip in under two seconds on CPU.
- New kokoro-inno-clone-tuner adapter adds zero-shot voice cloning to the frozen Kokoro-82M TTS model.
- Outputs standard Kokoro voice packs at shape [510, 1, 256] that drop into existing pipelines unchanged.
- Total model size around 24MB at fp16; enrollment takes 1.4s for a 30s reference on CPU.
- Trained on LibriTTS-R, VoxPopuli and Emilia-YODAS for under $20 in GPU time on HF Jobs.
- Hits SIM-o 0.288 on LibriSpeech test-clean, roughly double the nearest stock Kokoro pack.
- Already integrated into Kokoro-FastAPI v0.9.0+; Apache-2.0 licensed.
Kokoro-82M gains zero-shot voice tuning
Kokoro-82M combines fast synthesis, stable output, and a library of fixed voice packs. The new kokoro-inno-clone-tuner adapter adds zero-shot voice matching, which creates a voice from an unseen reference clip without per-speaker training. It keeps the base model frozen and produces native Kokoro voice packs, so existing pipelines can use the result without changing synthesis models.
A native voice pack from one clip
The adapter wraps Kokoro-82M and converts reference audio into a tensor with Kokoro’s standard [510, 1, 256] shape. Applications can pass that tensor directly to KPipeline or save it as a reusable voice file with torch.save.
Enroll and synthesize
# Shell
pip install inno-kokoroimport torch
from inno_kokoro.enroll import Tuner, enroll, read
from kokoro import KPipeline
tuner = Tuner()
pack, _ = enroll(*read("my_ref.wav"), tuner)
torch.save(pack, "my_voice.pt")
pipeline = KPipeline(lang_code="a")
result = next(
pipeline("Hello from a tuned voice.", voice=pack)
)
audio = result.audioReference clips must be between 3 and 30 seconds, contain one speaker, and have reasonably clean audio. The current release supports English.
Drop-in deployment, split licenses
Kokoro-FastAPI includes the tuner from version 0.9.0 onward, allowing self-hosted installations to add enrollment without rebuilding their serving layer. The bundled speaker encoder also avoids a separate model download.
- Adapter code and weights: Apache-2.0
- Bundled speaker encoder: CC BY-SA 3.0
Deployments that redistribute the full package should review the obligations of both licenses, particularly the attribution and share-alike terms attached to the speaker encoder.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.