Noctaluna Squeezes Alibaba's Noct-Q-Anime Into 8GB GPUs for Anime Art

A community fine-tune of Alibaba's compact Qwen-Image-2.1 delivers uncensored anime illustration in an INT8 build that runs on 12GB consumer GPUs.

·
·
Noctaluna Squeezes Alibaba's Noct-Q-Anime Into 8GB GPUs for Anime ArtPRO
Read2 min
TypeModel
  • Noctaluna released an uncensored anime fine-tune of Qwen-Image-2.1 with 6,000+ downloads.
  • The INT8 build weighs 7.3 GB and runs on 8-12 GB VRAM cards in ComfyUI.
  • Base model shrank from 20B/40.9GB to a 7B visual generator while adding native RGBA transparency.
  • Recommended settings: 25 steps, euler simple, CFG 3, prompts starting with "An anime illustration of..."
  • No trigger word needed; supports two-character scenes, explicit content, and long detailed prompts.
  • Locked to Qwen Research License, non-commercial use only, unlike the Apache 2.0 predecessors.

Noct-Q-Anime packages Qwen-Image-2.1 for 8–12GB GPUs

Noctaluna has published Noct-Q-Anime, a community derivative of Alibaba’s Qwen-Image-2.1 built for anime illustration. The checkpoint stores the visual-generator weights in INT8, merges in anime-focused transformer weights, and removes the base model’s content restrictions. Its Hugging Face listing recorded more than 6,000 downloads within the first few days.

The author targets ComfyUI systems with 8 to 12GB of VRAM. The diffusion checkpoint occupies 7.3GB, although peak memory also depends on image resolution, batch size, model offloading, and the separate text encoder and VAE.

Seven billion parameters cut the footprint

The checkpoint builds on Qwen-Image-2.1 notes, Alibaba’s latest downloadable image model. Alibaba reduced the visual generator from roughly 20 billion parameters in the previous Qwen-Image release to 7 billion, cutting its BF16 weight size from 40.9GB to 14.2GB.

Qwen-Image-2.1 component Specification
Visual generator 7 billion parameters
Architecture 32 single-stream diffusion transformer layers
BF16 weights 14.2GB
INT8 weights 7.26GB
Supported tasks Image generation, editing, and transparent output

A diffusion transformer, or DiT, progressively converts noise into an image representation. BF16 stores each weight with 16 bits, while INT8 uses 8-bit integers to reduce storage and memory requirements. Quantization can affect output quality, but it makes the visual generator practical on a wider range of consumer hardware.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads