YuE2 Now Generates Full Songs Locally Without Python or PyTorch
Pre-quantized GGUF weights and a portable C++ backend let YuE2 generate 48kHz stereo songs on a single consumer GPU with an editable ABC score.
- YuE2-GGUF ships pre-quantized weights (Q5_K_M to BF16) for local song generation.
- Runs via yue2.cpp, a C++17/GGML backend supporting CPU, CUDA, and Vulkan.
- AR half writes an editable ABC score, NAR half paints acoustic latents via flow matching.
- 65-second song at Q8_0 peaks at 5.8 GB VRAM, or 3.8 GB with reduced context.
- Includes optional SheetSage2 transcriber for audio-to-score covers of existing recordings.
- Weights are CC BY-NC 4.0, so no commercial use without upstream permission.
YuE2 brings local song generation to GGUF and C++
YuE2 can now generate complete songs through a native C++17 stack. The release combines YuE2 GGUF weights with the yue2.cpp runtime, built on GGML for CPU, CUDA, and Vulkan. Users provide style tags and lyrics, and the pipeline returns 48 kHz stereo audio alongside the ABC score it composed.
YuE2 plans before it renders
Multimodal Art Projection released YuE2-3B as an open-weight model for generating songs with vocals and accompaniment. Its authors report results competitive with Suno v5 and v6 on WildSongBench, although that claim is benchmark-specific and comes from the model team.
YuE2 first writes melody and chord information in ABC, a plain-text notation format, then generates the audio representation. That intermediate score gives developers an editable checkpoint between the prompt and the final track. The GGUF port preserves this process while replacing the Python and PyTorch runtime with native executables.
Three weight sets drive the pipeline
The release provides three model families converted from the upstream checkpoints:
| Component | Role | Available sizes |
|---|---|---|
| Backbone | A 3.6B-parameter Mixture-of-Transformers that generates the score, semantic codes, and acoustic latents. | 7.17 GB in BF16 to 2.62 GB in Q5_K_M. The download script selects the 3.81 GB Q8_0 build by default. |
| VAE | An Oobleck SnakeBeta decoder that converts acoustic latents into stereo audio. |
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.