Kokoro-82M Clones Any Voice From a 3-Second Clip for Under $20

A tiny 24MB adapter turns Kokoro-82M into a zero-shot voice cloner that enrolls a reference clip in under two seconds on CPU.

·
·
·
Kokoro-82M Clones Any Voice From a 3-Second Clip for Under $20PRO
  • New kokoro-inno-clone-tuner adapter adds zero-shot voice cloning to the frozen Kokoro-82M TTS model.
  • Outputs standard Kokoro voice packs at shape [510, 1, 256] that drop into existing pipelines unchanged.
  • Total model size around 24MB at fp16; enrollment takes 1.4s for a 30s reference on CPU.
  • Trained on LibriTTS-R, VoxPopuli and Emilia-YODAS for under $20 in GPU time on HF Jobs.
  • Hits SIM-o 0.288 on LibriSpeech test-clean, roughly double the nearest stock Kokoro pack.
  • Already integrated into Kokoro-FastAPI v0.9.0+; Apache-2.0 licensed.

Kokoro-82M gains zero-shot voice tuning

Kokoro-82M combines fast synthesis, stable output, and a library of fixed voice packs. The new kokoro-inno-clone-tuner adapter adds zero-shot voice matching, which creates a voice from an unseen reference clip without per-speaker training. It keeps the base model frozen and produces native Kokoro voice packs, so existing pipelines can use the result without changing synthesis models.

A native voice pack from one clip

The adapter wraps Kokoro-82M and converts reference audio into a tensor with Kokoro’s standard [510, 1, 256] shape. Applications can pass that tensor directly to KPipeline or save it as a reusable voice file with torch.save.

Enroll and synthesize

code
# Shell
pip install inno-kokoro
haskell
import torch

from inno_kokoro.enroll import Tuner, enroll, read
from kokoro import KPipeline

tuner = Tuner()
pack, _ = enroll(*read("my_ref.wav"), tuner)

torch.save(pack, "my_voice.pt")

pipeline = KPipeline(lang_code="a")
result = next(
    pipeline("Hello from a tuned voice.", voice=pack)
)
audio = result.audio

Reference clips must be between 3 and 30 seconds, contain one speaker, and have reasonably clean audio. The current release supports English.

Drop-in deployment, split licenses

Kokoro-FastAPI includes the tuner from version 0.9.0 onward, allowing self-hosted installations to add enrollment without rebuilding their serving layer. The bundled speaker encoder also avoids a separate model download.

  • Adapter code and weights: Apache-2.0
  • Bundled speaker encoder: CC BY-SA 3.0

Deployments that redistribute the full package should review the obligations of both licenses, particularly the attribution and share-alike terms attached to the speaker encoder.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads