ACE-Step 1.5 Lets Developers Generate Music Locally Without PyTorch

A C++ port with pre-quantized GGUF weights brings the ACE-Step 1.5 music foundation model to CPU, CUDA, Metal and Vulkan backends.

·
·
·
ACE-Step 1.5 Lets Developers Generate Music Locally Without PyTorchPRO
Read2 min
TypeModel
  • ACE-Step 1.5 is now available as pre-quantized GGUF for the acestep.cpp C++17 runtime.
  • Backends include CPU, CUDA, Metal and Vulkan, outputting stereo 48kHz audio from text and lyrics.
  • Pipeline splits into Qwen3 text encoder, Qwen3 LM (0.6B/1.7B/4B), flow-matching DiT (2B or 4B XL) and a BF16 VAE.
  • Quants range from BF16 to Q4_K_M, though 4B LM skips Q4_K_M because it breaks audio code generation.
  • Underlying model generates full songs in under 2s on A100 and under 4GB VRAM, per the ACE-Step 1.5 paper.
  • Supports cover generation, repainting, vocal-to-BGM conversion and 50+ languages, with LoRA personalization.

ACE-Step 1.5 gets GGUF weights and a C++ runtime

A community GGUF release packages the ACE-Step 1.5 music-generation components as quantized model files and pairs them with acestep.cpp, a portable C++17 runtime built on GGML. Developers can now generate 48 kHz stereo music locally without assembling a PyTorch environment, CUDA-specific wheels and full-precision checkpoints.

GGUF stores model tensors and metadata in a format designed for efficient loading and quantization. In this release, it brings the deployment pattern familiar from local language models to a text-to-music pipeline. The runtime accepts a text caption and optional lyrics, then runs on CPU, CUDA, Metal or Vulkan. Accelerator support and performance depend on the selected build flags, hardware and quantization.

What the release contains

Area Details
Distribution Community GGUF conversions plus a dedicated C++ runtime
Input Text caption, optional lyrics and generation settings
Output Stereo 48 kHz audio
Interfaces Web server and separate LM and synthesis command-line tools
Backends CPU, CUDA, Metal and Vulkan
Default download About 7.7 GB for the Q8_0 turbo model set

Four files, two generation stages

ACE-Step separates planning, audio-code generation and waveform synthesis across several models. That design lets developers choose the language model and diffusion transformer independently, while the text encoder and VAE remain tied to the training architecture.

Component Role Available variants
Text encoder Converts captions and lyrics into conditioning vectors used by the diffusion transformer. Replacing it with another embedding model would break the representation learned during training. Qwen3-Embedding-0.6B in BF16 or Q8_0
Language model Expands the request into song metadata, lyrics and a 5 Hz audio-code sequence. At 5 Hz, each code covers 200 milliseconds; the learned code vocabulary contains 64,000 entries. Qwen3-based 0.6B, 1.7B and 4B models
Diffusion transformer Uses flow matching to turn the plan into 25 Hz audio latents, adding timbre, transients and stereo structure.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads