Unsloth Shrinks Alibaba's Z-Image-Turbo to 3.64 GB for Consumer GPUs
Unsloth released GGUF quantizations of Tongyi's Z-Image-Turbo, a 6B parameter text-to-image model that runs in under 16GB VRAM with 8-step inference.
- Unsloth released GGUF quantizations of Z-Image-Turbo, from 3.64 GB (Q2_K) to 12.3 GB (BF16).
- Base model is a 6B parameter Single-Stream DiT from Alibaba Tongyi-MAI, Apache 2.0 licensed.
- Generates 1024x1024 images in 8 function evaluations, sub-second on H800, fits in 16 GB consumer VRAM.
- Claims SOTA among open-source models on Alibaba AI Arena Elo leaderboard for text-to-image.
- Strong bilingual (English and Chinese) text rendering, a weak spot for most open models.
- Powered by new Decoupled-DMD distillation and DMDR RL post-training recipes.
Z-Image-Turbo shrinks a 6B image model to 3.64 GB
Unsloth has released quantized GGUF builds of Alibaba Tongyi-MAI’s Z-Image-Turbo, a distilled diffusion transformer that generates images with eight function evaluations. Available files range from a 3.64 GB Q2_K build to 12.3 GB BF16 and F16 weights.
Developers can trade precision for a smaller memory footprint, bringing the model within reach of more consumer GPUs through quantization and CPU offloading. The model also offers Apache 2.0 licensing, English and Chinese text rendering, and reported sub-second generation on an H800. At publication, the GGUF repository displayed more than 800,000 downloads. Hugging Face’s counter tracks file downloads and can include multiple downloads per user.
- Model size: 6 billion parameters
- Generation schedule: Eight function evaluations, exposed as nine pipeline steps
- Smallest GGUF: 3.64 GB in Q2_K
- License: Apache 2.0
- Pending releases: Z-Image-Base and Z-Image-Edit
One stream carries every modality
Alibaba built Z-Image around a Scalable Single-Stream Diffusion Transformer, abbreviated S3-DiT. The architecture places text tokens, visual semantic tokens, and compressed VAE image tokens in one sequence, allowing the same transformer layers to process information from every modality.
Dual-stream architectures such as SD3 and FLUX allocate separate transformer branches to text and image representations before combining them. Z-Image’s shared stream aims to use its parameter budget more efficiently. The VAE supplies a compact representation of the image, while visual semantic tokens carry higher-level information about its contents.
Turbo targets eight evaluations
Z-Image-Turbo is the distilled member of the model family. Its recommended Diffusers configuration uses num_inference_steps=9 and guidance_scale=0.0, corresponding to the project’s eight-function-evaluation generation recipe.
Tongyi-MAI reports sub-second latency on an enterprise H800 GPU and targets operation within 16 GB of VRAM on consumer hardware. Actual latency and memory use depend on resolution, compute precision, attention implementation, offloading, and whether the text encoder and VAE remain on the GPU.
The planned Z-Image-Base checkpoint will provide an undistilled foundation for fine-tuning, while Z-Image-Edit will add image-to-image generation. Neither checkpoint was available with the Turbo release.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.