Mitsuba Squeezes a 27B Vision Model Into 7.3 GB for ComfyUI
A 7.3GB ternary-quantized Qwen3.8-27B vision model tuned specifically for writing ComfyUI image and video generation prompts from reference images.
- Mitsuba-ComfyUI-27B is a ternary-quantized Qwen3.8-27B VLM, just 7.3GB.
- Specialized for image-to-prompt workflows in ComfyUI, Krea 2, and similar pipelines.
- Vision score barely drops (89.8 to 87.8); prompt generation success actually improves (3/10 to 6/10).
- Coding ability collapses to 4/100; not recommended for anything but prompt and caption work.
- Requires the PrismML llama.cpp fork and reasoning mode turned off.
- Apache 2.0 license, runs on a single 16GB GPU at roughly 119 tokens/sec on an RTX 5090.
Mitsuba compresses a 27B vision model for ComfyUI prompts
Mitsuba-ComfyUI-27B-GGUF is a 27-billion-parameter vision-language model that analyzes images and writes constrained prompts for downstream image and video generators. The Japanese developer publishing as isichan-ai tuned the Qwen3.8-27B derivative for Stable Diffusion tags, timed video instructions, negative prompts and detailed image descriptions.
Within a ComfyUI graph, Mitsuba can receive an image and prompt requirements, then return text for a diffusion or video-generation node. This gives local workflows an image-aware prompt-writing stage without relying on an external captioning or prompt service. Coding falls outside the model’s intended scope.
A 27B model squeezed below 8 GB
| Component or measurement | Reported value |
|---|---|
| Language backbone | Qwen3.8-27B |
| Recommended weights | PQ2_0, 7.3 GB |
| Smaller weights | PTQ1_0, 6.0 GB |
| Vision component | mmproj-Q8_0.gguf, 0.63 GB |
| Decode throughput | 119 tokens per second with PQ2_0 on an RTX 5090 |
| GPU target | One 16 GB GPU, according to the author |
| License | Apache 2.0 |
Ternary quantization restricts each compressed weight to three values, approximately -1, 0 and +1. Encoding three states requires about 1.58 bits in theory, which explains the model’s unusually small GGUF files. GGUF is the model format commonly loaded by llama.cpp-based runtimes.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.