ACE-Step 1.5 Lets Developers Generate Music Locally Without PyTorch
A C++ port with pre-quantized GGUF weights brings the ACE-Step 1.5 music foundation model to CPU, CUDA, Metal and Vulkan backends.
- ACE-Step 1.5 is now available as pre-quantized GGUF for the acestep.cpp C++17 runtime.
- Backends include CPU, CUDA, Metal and Vulkan, outputting stereo 48kHz audio from text and lyrics.
- Pipeline splits into Qwen3 text encoder, Qwen3 LM (0.6B/1.7B/4B), flow-matching DiT (2B or 4B XL) and a BF16 VAE.
- Quants range from BF16 to Q4_K_M, though 4B LM skips Q4_K_M because it breaks audio code generation.
- Underlying model generates full songs in under 2s on A100 and under 4GB VRAM, per the ACE-Step 1.5 paper.
- Supports cover generation, repainting, vocal-to-BGM conversion and 50+ languages, with LoRA personalization.
ACE-Step 1.5 gets GGUF weights and a C++ runtime
A community GGUF release packages the ACE-Step 1.5 music-generation components as quantized model files and pairs them with acestep.cpp, a portable C++17 runtime built on GGML. Developers can now generate 48 kHz stereo music locally without assembling a PyTorch environment, CUDA-specific wheels and full-precision checkpoints.
GGUF stores model tensors and metadata in a format designed for efficient loading and quantization. In this release, it brings the deployment pattern familiar from local language models to a text-to-music pipeline. The runtime accepts a text caption and optional lyrics, then runs on CPU, CUDA, Metal or Vulkan. Accelerator support and performance depend on the selected build flags, hardware and quantization.
What the release contains
| Area | Details |
|---|---|
| Distribution | Community GGUF conversions plus a dedicated C++ runtime |
| Input | Text caption, optional lyrics and generation settings |
| Output | Stereo 48 kHz audio |
| Interfaces | Web server and separate LM and synthesis command-line tools |
| Backends | CPU, CUDA, Metal and Vulkan |
| Default download | About 7.7 GB for the Q8_0 turbo model set |
Four files, two generation stages
ACE-Step separates planning, audio-code generation and waveform synthesis across several models. That design lets developers choose the language model and diffusion transformer independently, while the text encoder and VAE remain tied to the training architecture.
| Component | Role | Available variants |
|---|---|---|
| Text encoder | Converts captions and lyrics into conditioning vectors used by the diffusion transformer. Replacing it with another embedding model would break the representation learned during training. | Qwen3-Embedding-0.6B in BF16 or Q8_0 |
| Language model | Expands the request into song metadata, lyrics and a 5 Hz audio-code sequence. At 5 Hz, each code covers 200 milliseconds; the learned code vocabulary contains 64,000 entries. | Qwen3-based 0.6B, 1.7B and 4B models |
| Diffusion transformer | Uses flow matching to turn the plan into 25 Hz audio latents, adding timbre, transients and stereo structure. |
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.