Mia's AI Lab Squeezes GLM-5.3-Flash Closer to Full Precision at 4 Bits

A recalibrated 4-bit EXL3 quant of GLM-5.3-Flash matches the original model closer than the standard TR3 baseline at identical size and speed.

·
·
·
Mia's AI Lab Squeezes GLM-5.3-Flash Closer to Full Precision at 4 BitsPRO
Read2 min
TypeModel
TopicGpus · Llms
  • Mia's AI Lab released a new EXL3 4bpw quant of GLM-5.3-Flash tuned for TensorFold on DGX Spark.
  • KL divergence to the original model drops 2-18% as served, 11-27% experts-only, across six test sets.
  • Confident-token mistakes drop 16-37% versus the TR3-4bpw baseline at the same size and throughput.
  • HumanEval+/MBPP+ pass rate is statistically tied, but replies are about 10% shorter at the same accuracy.
  • v1.4 adds a TP=3 path for 3x DGX Sparks: 77 tok/s single-stream, 146 tok/s at 4 streams, ~6M KV cache.
  • Only routed experts are quantized; dense layers stay BF16 and are quantized at load by TensorFold.

GLM-5.3-Flash Gets a TensorFold-Tuned 4-Bit Quant

Mia’s AI Lab has released an EXL3 4-bit quantization of Z.AI’s GLM-5.3-Flash, tuned for the TensorFold serving engine on NVIDIA DGX Spark hardware. At the same bit width, file size, and reported throughput as the TR3-4bpw baseline, the new checkpoint produces next-token probabilities closer to the full-precision model across the team’s tests.

The Apache 2.0 Hugging Face release includes a serving recipe for two DGX Spark systems. Downloading requires acceptance of the model page’s gated-access terms. All quality and performance results below come from Mia’s AI Lab and have not been independently verified.

Four bits, less drift

KL divergence measures how far a compressed model’s next-token probability distribution moves from its full-precision reference; lower values indicate a closer match. Mia reports lower point estimates on all eight test sets. The reduction ranges from 11% to 27% when measuring the quantized routed experts alone, and from 2% to 18% under TensorFold’s serving configuration, which also converts the dense weights to 4-bit precision.

Paired KL divergence changes with 95% confidence intervals across eight test sets
Paired KL divergence changes relative to TR3-4bpw. Lower values favor the new quantization.

Both checkpoints were scored on identical tokens, allowing a paired statistical comparison. Seven of the eight 95% confidence intervals fall entirely below zero; the chat set crosses zero, leaving that result inconclusive. On tokens where the full-precision model assigned high confidence to its top choice, the experts-only quantization selected a different token 16% to 37% less often than TR3.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads