Alibaba's Qwen-Image-2.1 Tops Two Open-Weight Image Leaderboards
Alibaba's 7B unified generation and editing model sweeps 16 of 36 categories on Artificial Analysis leaderboards, dethroning Ideogram and HunyuanImage among open weights.
- Qwen-Image-2.1 is now the #1 open weights model on both Artificial Analysis T2I and editing leaderboards.
- Leads open weights in 16 of 36 total category leaderboards across capabilities and use cases.
- 7B unified generation plus editing, native 2K output, native RGBA transparency, up to 10 reference images.
- Architecture: 32-layer single-stream DiT, Qwen3-VL 8B text encoder, 64-channel RGBA VAE.
- Released under Qwen Research License, non-commercial only, commercial use needs separate agreement.
- Day-zero support in Diffusers, ComfyUI, vLLM-Omni, SGLang and LightX2V; 33 GB download on Hugging Face.
Qwen-Image-2.1 leads two open-weight image leaderboards
Alibaba’s Qwen-Image-2.1 now ranks No. 1 among open-weight models on the Artificial Analysis AA-Image-T2I v2.0 and AA-Image-Editing v2.0 leaderboards. It displaced Ideogram 4.0 (Quality) in text-to-image generation and HunyuanImage 3.0 Instruct in image editing.
For developers, the release combines generation, editing, transparent output, and multi-image conditioning in one downloadable checkpoint. Open-weight describes access to the model parameters. Deployment rights come from a separate research license that restricts commercial use.
Across fields that include proprietary systems, Qwen-Image-2.1 ranks 18th for both text-to-image generation and editing on Artificial Analysis. The previous Qwen Image 2.0 ranked 72nd and 58th, respectively. Qwen-Image-2.1 also leads 16 of the benchmark’s 36 capability and use-case categories.
One checkpoint, two image jobs
Alibaba’s launch post describes a single checkpoint for text-to-image generation and instruction-based editing. It produces native 2,048 × 2,048 output, accepts as many as 10 reference images per generation, and supports RGBA output with transparency.
| Component | Specification |
|---|---|
| Image backbone | 32-layer, single-stream diffusion transformer with 7 billion parameters |
| Prompt encoder | Qwen3-VL 8B |
| Image codec | 64-channel RGBA variational autoencoder with 16× spatial compression |
| Native output | 2,048 × 2,048 pixels |
| Reference input | Up to 10 images per generation |
| Download size | Approximately 33 GB across the model components |
| License | Qwen Research License Agreement, with commercial use requiring separate permission |
A diffusion transformer, commonly shortened to DiT, applies a transformer architecture to the denoising process used to generate images. The Qwen3-VL encoder interprets prompts and reference inputs, while the variational autoencoder compresses images into a smaller latent representation and decodes the result.
The RGBA autoencoder preserves an alpha channel alongside red, green, and blue. Compatible pipelines can therefore save transparent PNG assets directly, avoiding a separate background-removal stage.
The complete download spans the image backbone, encoder, and autoencoder. Runtime memory depends on precision, resolution, quantization, component offloading, and whether the pipeline keeps every component loaded on the GPU.
- Weights: Available through Hugging Face and ModelScope.
- Python pipelines: Diffusers added release-day support.
- Visual workflows: ComfyUI provides a ready-made template.
- Serving stacks: vLLM-Omni, SGLang, and LightX2V support the model.
Sixteen category leads
Artificial Analysis divides its evaluation into nine text-to-image capabilities, 10 text-to-image use cases, seven editing actions, and 10 editing use cases. Qwen-Image-2.1 leads the open-weight field in the following categories:
- Text-to-image capabilities, 5 of 9: Complex Compositions, Text Rendering, Knowledge, Layout, and Lighting.
- Text-to-image use cases, 3 of 10: Social Media and Creator Content, Productivity and Knowledge Work, and Consumer. It also scores level with the leader in Architecture and Frontier.
- Editing actions, 3 of 7: Text or Symbol Edits, Reasoning-Based Edit, and Composition and Framing.
- Editing use cases, 5 of 10: Productivity and Knowledge Work, Social Media and Creator Content, Retail and Ecommerce, Animation and Gaming, and UI/UX Design.
| Leaderboard | Qwen-Image-2.1 | Next model | Margin |
|---|---|---|---|
| AA-Image-Editing v2.0 | 1,074 | HunyuanImage 3.0 Instruct, 1,066 | +8 |
| AA-Image-T2I v2.0 | 1,034 | Ideogram 4.0 (Quality), 1,011 | +23 |
Elo converts pairwise preference results into a relative rating. Scores are comparable within the same leaderboard and can change as new evaluations arrive. Qwen Image Edit Plus 2511 follows the two editing leaders with an Elo rating of 1,021.
Where the scores peak
Relative to the best results across open and proprietary models, Qwen-Image-2.1 performs most strongly in Lighting, Knowledge, and Reasoning. Those categories test details such as reflections, refractions, shadows, landmarks, species, domain facts, spatial relationships, mathematical concepts, and mixed visual ideas.
Compared with Qwen Image 2.0, the new checkpoint narrows the measured gap across every text-to-image capability. Its largest gains appear in Layout, Lighting, and Knowledge.
For editing, Qwen-Image-2.1 comes closest to the overall leader in Scene and Style Edit, which includes relighting, restyling, and background replacement. Text or Symbol Edits and Enhancement and Restoration follow. The largest generational gains appear in Text or Symbol Edits, Object-Level Edit, and Identity-Preserving Edit.
Closed systems continue to lead the overall field. On LMArena, Qwen-Image-2.1 ranks 16th for editing and 17th for text-to-image generation. Models from OpenAI, Google, Microsoft, Meta, xAI, and other vendors occupy most higher positions. Its first-place claim applies specifically to open-weight models on the Artificial Analysis boards.
Commercial terms narrow deployment
The Qwen Research License Agreement grants use “for non-commercial purposes only.” Commercial deployment requires a separate license requested through Alibaba’s licensing email.
Earlier Qwen-Image releases used Apache 2.0, which permits commercial use subject to its terms. Teams moving to version 2.1 need to account for the new agreement before integrating the checkpoint into a paid product, hosted service, or revenue-generating workflow.
Hardware and workflow fit
The 7-billion-parameter image backbone may run on a high-memory consumer GPU with reduced precision, quantization, CPU offloading, or staged component loading. The approximately 33 GB package exceeds the VRAM available on many consumer cards, while native 2K generation adds further memory pressure. Exact requirements depend on the chosen runtime and optimization settings.
The benchmark profile and feature set align most closely with these workloads:
- Transparent assets: Sprites, icons, interface panels, stickers, and product cutouts with an alpha channel.
- Information graphics: Diagrams, infographics, charts, and slides that depend on layout and text rendering.
- Promotional graphics: Thumbnails, social posts, and cards containing prominent text.
- Multi-reference composition: Workflows combining characters, products, backgrounds, and style references.
- Instruction-based editing: Changes that require spatial reasoning, identity preservation, or interpretation of a complex request.
Developers can compare outputs in the Image Arena, download the weights from Hugging Face or ModelScope, or use the supported ComfyUI workflow. Noncommercial self-hosting requires suitable hardware and acceptance of the research license; shipping a commercial product requires Alibaba’s approval.