Alibaba's Qwen-Image-2.1 Now Runs on Apple Silicon at Just 7 GB

A 4-bit MLX port of Alibaba's 7B image model runs on Apple Silicon with 7GB peak memory, keeping generation and editing local.

·
·
Alibaba's Qwen-Image-2.1 Now Runs on Apple Silicon at Just 7 GBPRO
  • JoyFusionAI released a 4-bit MLX build of Qwen-Image-2.1 for Apple Silicon via mflux.
  • Peak memory drops to about 7 GB on M1 Max at 768x768 with the low-ram flag.
  • Roughly 6.7 seconds per step, 40 steps recommended, no classifier-free guidance needed.
  • Upstream model unifies text-to-image, editing, and native RGBA transparency in one 7B DiT.
  • Accepts up to 10 reference images and local edits via circles, brush, or mask.
  • Weights sit under the Qwen Research License, so commercial use requires a separate agreement.

Qwen-Image-2.1 Gets a 4-Bit MLX Build for Apple Silicon

Alibaba’s Qwen-Image-2.1 is now available as a JoyFusionAI checkpoint for Apple Silicon. The package uses MLX, Apple’s machine-learning framework for unified memory, and runs through the mflux image-generation pipeline. Quantization reduces the model’s generation transformer to 3.8 GiB, making local generation and editing practical on Macs with limited memory.

Qwen-Image-2.1 combines text-to-image generation, reference-based editing, and transparent RGBA output in one model. Its visual generator has 7 billion parameters, about one-third of the original Qwen-Image model’s 20 billion. The August 2025 release also required a separate Qwen-Image-Edit checkpoint, while version 2.1 handles both tasks. Qwen reports a 60.28 composite score in its published comparison and says the model leads alternatives with downloadable weights.

A 20 GB Package With a 3.8 GiB Transformer

JoyFusionAI created the checkpoint from the upstream repository with mflux-save -q 4. The 4-bit label applies primarily to the diffusion transformer, the network that progressively turns noise into an image. The text encoder remains in bf16, so it accounts for most of the download.

Component Format Size
7B single-stream DiT 4-bit MLX affine, group size 64 3.8 GiB
Qwen3-VL 8B text encoder bf16 14 GiB
64-channel VAE 4-bit linear layers and

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads