Mitsuba Squeezes a 27B Vision Model Into 7.3 GB for ComfyUI

A 7.3GB ternary-quantized Qwen3.8-27B vision model tuned specifically for writing ComfyUI image and video generation prompts from reference images.

·
·
Mitsuba Squeezes a 27B Vision Model Into 7.3 GB for ComfyUIPRO
  • Mitsuba-ComfyUI-27B is a ternary-quantized Qwen3.8-27B VLM, just 7.3GB.
  • Specialized for image-to-prompt workflows in ComfyUI, Krea 2, and similar pipelines.
  • Vision score barely drops (89.8 to 87.8); prompt generation success actually improves (3/10 to 6/10).
  • Coding ability collapses to 4/100; not recommended for anything but prompt and caption work.
  • Requires the PrismML llama.cpp fork and reasoning mode turned off.
  • Apache 2.0 license, runs on a single 16GB GPU at roughly 119 tokens/sec on an RTX 5090.

Mitsuba compresses a 27B vision model for ComfyUI prompts

Mitsuba-ComfyUI-27B-GGUF is a 27-billion-parameter vision-language model that analyzes images and writes constrained prompts for downstream image and video generators. The Japanese developer publishing as isichan-ai tuned the Qwen3.8-27B derivative for Stable Diffusion tags, timed video instructions, negative prompts and detailed image descriptions.

Within a ComfyUI graph, Mitsuba can receive an image and prompt requirements, then return text for a diffusion or video-generation node. This gives local workflows an image-aware prompt-writing stage without relying on an external captioning or prompt service. Coding falls outside the model’s intended scope.

A 27B model squeezed below 8 GB

Release specifications reported on the model card
Component or measurement Reported value
Language backbone Qwen3.8-27B
Recommended weights PQ2_0, 7.3 GB
Smaller weights PTQ1_0, 6.0 GB
Vision component mmproj-Q8_0.gguf, 0.63 GB
Decode throughput 119 tokens per second with PQ2_0 on an RTX 5090
GPU target One 16 GB GPU, according to the author
License Apache 2.0

Ternary quantization restricts each compressed weight to three values, approximately -1, 0 and +1. Encoding three states requires about 1.58 bits in theory, which explains the model’s unusually small GGUF files. GGUF is the model format commonly loaded by llama.cpp-based runtimes.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads