ByteShape Squeezes Qwen3.8-27B Into 8.8 GB Hitting 176 Tokens per Second
ByteShape ships ShapeLearn-quantized GGUFs of Qwen3.8-27B that fit on 12 GB GPUs and hit 176 tokens per second on an RTX 5090.
- ByteShape released ShapeLearn-quantized Qwen3.8-27B GGUFs in five sizes from 2.56 to 3.84 bits per weight.
- The smallest build fits in 8.8 GB VRAM; the largest IQ4_XS variant is 13.1 GB.
- Per-tensor quantization is scored against BF16 on GSM8K, MMLU, LiveCodeBench, IFEval, BFCL and ACEBench.
- Embedded MTP head adds less than 250 MB and delivers 1.28 to 1.66x decode speedup.
- External DFlash 2 draft reaches 1.34 to 2.10x speedup, text-only, needs llama.cpp b10658+.
- Apache 2.0 license, vision capable, runs via llama.cpp, Ollama, LM Studio, vLLM and SGLang.
ByteShape compresses Qwen3.8-27B below 14 GB
ByteShape has released five GGUF quantizations of the Qwen3.8-27B vision-language model. The 8.8 GB to 13.1 GB files use ShapeLearn, a per-tensor quantization system designed to preserve sensitive weights at low bit depths. In ByteShape’s RTX 5090 tests, speculative decoding reached 176 tokens per second with the smallest checkpoint.
A 27-billion-parameter model stored in BF16 requires roughly 54 GB for weights before runtime overhead. Quantization cuts that requirement by storing weights with fewer bits. ShapeLearn selects the datatype and quantization method for each tensor, assigning more precision to weights its optimizer identifies as sensitive. GGUF packages those quantized weights for llama.cpp-compatible local runtimes.
Five checkpoints below 14 GB
The five GPU-oriented builds span 2.56 to 3.84 average bits per weight, abbreviated as bpw. Lower values reduce storage and memory use, usually at the cost of some model quality.
| Checkpoint | Bits per weight | File size |
|---|---|---|
| GPU-1 | 2.56 bpw | 8.8 GB |
| GPU-2 | 2.88 bpw | 9.9 GB |
| GPU-3 | 3.01 bpw | 10.4 GB |
| GPU-4 | 3.23 bpw | 11.0 GB |
| GPU-5 | 3.84 bpw | 13.1 GB |
A 16 GB card has enough capacity for each checkpoint file, although total VRAM use also includes the KV cache, runtime buffers, vision projector, and any speculative draft model. Context length, cache format, and GPU offloading determine whether a particular configuration fits entirely in memory.
Filename tags such as IQ4_XS and IQ2_XXS provide compatible indexing on Hugging Face. The underlying files contain ShapeLearn’s hybrid mix of quantization methods, so the tag alone does not describe every tensor’s format.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.