Mia's AI Lab Squeezes GLM-5.3-Flash Closer to Full Precision at 4 Bits
A recalibrated 4-bit EXL3 quant of GLM-5.3-Flash matches the original model closer than the standard TR3 baseline at identical size and speed.
- Mia's AI Lab released a new EXL3 4bpw quant of GLM-5.3-Flash tuned for TensorFold on DGX Spark.
- KL divergence to the original model drops 2-18% as served, 11-27% experts-only, across six test sets.
- Confident-token mistakes drop 16-37% versus the TR3-4bpw baseline at the same size and throughput.
- HumanEval+/MBPP+ pass rate is statistically tied, but replies are about 10% shorter at the same accuracy.
- v1.4 adds a TP=3 path for 3x DGX Sparks: 77 tok/s single-stream, 146 tok/s at 4 streams, ~6M KV cache.
- Only routed experts are quantized; dense layers stay BF16 and are quantized at load by TensorFold.
GLM-5.3-Flash Gets a TensorFold-Tuned 4-Bit Quant
Mia’s AI Lab has released an EXL3 4-bit quantization of Z.AI’s GLM-5.3-Flash, tuned for the TensorFold serving engine on NVIDIA DGX Spark hardware. At the same bit width, file size, and reported throughput as the TR3-4bpw baseline, the new checkpoint produces next-token probabilities closer to the full-precision model across the team’s tests.
The Apache 2.0 Hugging Face release includes a serving recipe for two DGX Spark systems. Downloading requires acceptance of the model page’s gated-access terms. All quality and performance results below come from Mia’s AI Lab and have not been independently verified.
Four bits, less drift
KL divergence measures how far a compressed model’s next-token probability distribution moves from its full-precision reference; lower values indicate a closer match. Mia reports lower point estimates on all eight test sets. The reduction ranges from 11% to 27% when measuring the quantized routed experts alone, and from 2% to 18% under TensorFold’s serving configuration, which also converts the dense weights to 4-bit precision.
Both checkpoints were scored on identical tokens, allowing a paired statistical comparison. Seven of the eight 95% confidence intervals fall entirely below zero; the chat set crosses zero, leaving that result inconclusive. On tokens where the full-precision model assigned high confidence to its top choice, the experts-only quantization selected a different token 16% to 37% less often than TR3.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.