Open-Source VieNeu-TTS v3 Turbo Runs 48 kHz Vietnamese Speech on CPUs

A from-scratch Vietnamese TTS model ships 48 kHz speech, 23 preset voices, instant cloning, and 16 concurrent real-time streams on a single RTX 3060.

·
·
·
Open-Source VieNeu-TTS v3 Turbo Runs 48 kHz Vietnamese Speech on CPUsPRO
  • VieNeu-TTS v3 Turbo ships 48 kHz Vietnamese TTS, Apache-2.0, trained from scratch on ~10k hours.
  • 16 concurrent real-time streams on a single RTX 3060 with ~115 ms first-audio latency.
  • 23 preset voices across Northern, Central, and Southern dialects, plus instant voice cloning.
  • OpenAI-compatible /v1/audio/speech endpoint drops into Pipecat, LiveKit, and OpenAI SDK clients.
  • CPU path is torch-free via ONNX Runtime; GPU path auto-switches to PyTorch with batching.
  • Supports English-Vietnamese code-switching, inline emotion cues, and single-GPU LoRA fine-tuning.

VieNeu-TTS v3 Turbo offers 48 kHz Vietnamese speech synthesis on CPUs

VieNeu-TTS v3 Turbo is an open-source text-to-speech model with 48 kHz output, voice cloning, Vietnamese-English code-switching, inline emotion cues, and an OpenAI-compatible streaming API. Its Hugging Face page reports nearly 780,000 downloads.

The project describes v3 Turbo as an original architecture trained from scratch on approximately 10,000 hours of English and Vietnamese speech. At roughly 100 million parameters, the model supports CPU inference while providing a separate CUDA path for higher concurrency. The weights use the Apache-2.0 license.

One checkpoint, two runtimes

The release publishes a single checkpoint in ONNX, Safetensors, and GGUF formats, along with a Python SDK named vieneu. CPU inference runs through ONNX Runtime without importing PyTorch. On CUDA systems, the SDK selects a PyTorch engine with automatic batching and a continuous-batching scheduler. Applications use the same API on either backend.

Voices, cloning, and control

  • 48 kHz audio: The model uses the MOSS Audio Tokenizer neural codec, doubling the 24 kHz output rate of the v2 series.
  • 23 preset voices: The 23 preset voices cover Northern, Central, and Southern Vietnamese regions with several reading characters.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads