NVIDIA's BigVGAN v2 Hits 44.1 kHz Audio With 3x Faster CUDA Inference
NVIDIA's BigVGAN v2 vocoder now runs at 44kHz with 512x upsampling, a fused CUDA kernel for 1.5-3x faster synthesis, and training across diverse audio.
- NVIDIA released BigVGAN v2, a 122M-parameter neural vocoder at 44kHz with 512x upsampling.
- Custom fused CUDA kernel delivers 1.5 to 3x faster inference on a single A100.
- Trained with multi-scale sub-band CQT discriminator and multi-scale mel spectrogram loss for 5M steps.
- Training data spans multilingual speech, environmental sounds, and musical instruments, not just English speech.
- Hugging Face integration lets you load and run with
from_pretrainedin a few lines. - Already embedded in 100+ Spaces including Seed-VC and MMAudio pipelines.
BigVGAN v2 adds 44.1 kHz vocoding and faster CUDA inference
NVIDIA’s BigVGAN v2 release adds pretrained neural vocoders for 22.05, 24, and 44.1 kHz audio. The flagship checkpoint converts 128-bin mel spectrograms into 44.1 kHz waveforms, using 512 waveform samples for each mel frame. A custom CUDA extension accelerates the anti-aliased activation blocks that perform much of the generator’s work.
A neural vocoder handles the final stage of many speech, music, voice-conversion, and video-to-audio pipelines. It receives a compact time-frequency representation called a mel spectrogram and synthesizes the raw samples sent to an audio file or playback device. Its speed and accuracy directly affect end-to-end latency, high-frequency detail, and audible artifacts.
A wider training set meets a fused kernel
- Five v2 configurations are provided in the pretrained models. The published checkpoints cover 22, 24, and 44 kHz output with 80, 100, or 128 mel bands.
- Fused CUDA inference. NVIDIA reports a 1.5x to 3x speedup over the standard PyTorch path on a single A100 GPU.
- Reworked objectives. Training adds a multi-scale sub-band constant-Q transform discriminator and a multi-scale mel-spectrogram loss.
- Broader source audio. NVIDIA describes a large compilation spanning multilingual speech, music, instruments, and environmental sounds.
The 44.1 kHz model with 512x upsampling contains about 122 million parameters and was trained for five million steps. The other v2 checkpoints contain about 112 million parameters. NVIDIA’s speed figure measures the generator’s fused CUDA path; complete pipeline throughput also depends on mel generation, data transfer, batching, and audio encoding.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.