Mitsuba Squeezes a 27B Vision Model Into 7.3 GB on One GPU
A ternary-quantized 27B vision-language model shrinks to 7.3 GB and specializes in writing ComfyUI image and video generation prompts.
- Mitsuba is a ternary 1.58-bit quantization of Qwen3.8-27B, shrunk to 7.3 GB for a single 16 GB GPU.
- Purpose-built for ComfyUI: image to prompt generation for Stable Diffusion, Krea, and video pipelines.
- Vision score 87.8 versus 89.8 for the BF16 original; coding collapsed from 66 to 4.
- Optional 70 MB HiMitsuba LoRA cuts both outright refusals and evasive answers on sensitive requests.
- Requires the PrismML llama.cpp fork; upstream does not yet support PQ2_0 or PTQ1_0.
- Must be run with thinking mode OFF or responses frequently come back empty.
Mitsuba squeezes a 27B vision model into 7.3 GB
Mitsuba is a ternary-quantized vision-language model designed to convert reference images into structured prompts for ComfyUI, Stable Diffusion, Krea, and similar generation tools. Its 27-billion-parameter base compresses to 7.3 GB or 6.0 GB, depending on the format, and the author reports that either build can run on a single 16 GB GPU.
The release targets developers who need local image analysis and tightly constrained prompt writing. General chat, code generation, and long-document analysis fall outside its strengths.
Image in, diffusion prompt out
Mitsuba accepts an image and describes it in a form that another model can use for image or video generation. Its supported workflows include:
- Stable Diffusion prompts that follow word limits, include required terms, exclude forbidden terms, and end with a
Negativeline. - Video prompts divided into intervals such as
0-3s,3-6s, and6-9s, with a camera movement assigned to each segment. - Image descriptions covering objects, quantities, people, scenes, charts, and visible text.
Prompt generation remains imperfect under strict constraints. The model card reports that Mitsuba satisfied every condition in six of ten test cases, so applications should validate required words, exclusions, length limits, and output structure before passing a response downstream.
Three-value weights, two builds
Mitsuba quantizes the official Qwen3.8-27B weights to roughly 1.58 bits per weight. Each weight is represented by one of three values: -1, 0, or +1. The files use Prism ML’s GGUF quantization formats.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.