Infatoshi Squeezes GLM-5.3's 753B Parameters Into 273 GiB for Multi-GPU Workstations
A 3-bit EXL3 quant squeezes a 753B uncensored GLM-5.3 Mixture of Experts into 273 GiB, making local inference possible on four RTX PRO 6000s.
- Infatoshi released a 3.0bpw EXL3 quant of GLM-5.3-UNCENSORED, totaling 273 GiB.
- Base is dealignai's weight-edited GLM-5.3, not a fine-tune, baked into residual-writer tensors.
- Architecture: 753B parameters, 256 routed experts (8 active), MLA attention with DSA sparse indexer, plus MTP draft layer.
- KL divergence vs FP8 source is 0.089; perplexity drifts from 3.302 to 3.440 on wikitext-2.
- Runs on 4x RTX PRO 6000s (384 GB VRAM) at roughly 55-60 tok/s without speculative decoding.
- Reports 0% refusals on HarmBench-320 at max effort, with practical 131K context ceiling on TP8 H200.
GLM-5.3’s 753B weights shrink to 273 GiB
Infatoshi has published an EXL3 release of GLM-5.3 that averages 3.04 bits per weight and occupies 273 GiB. The compression makes the 753-billion-parameter Mixture-of-Experts model practical on a high-end, multi-GPU workstation, although it remains far beyond a single consumer GPU.
The artifact quantizes dealignai’s GLM-5.3-UNCENSORED-FP8, a weight-edited version of zai-org’s original GLM-5.3. Developers evaluating it should account for both changes: the parent modifies refusal behavior, while the EXL3 conversion reduces numerical precision.
A mixed-precision squeeze
| Item | Detail |
|---|---|
| Architecture | GlmMoeDsaForCausalLM |
| Total parameters | 753 billion |
| Routed experts | 256, with 8 active per token |
| Shared experts | 1 |
| Layers | 78 transformer layers and 1 MTP layer |
| Attention | MLA with a DSA sparse indexer |
| Quantization | EXL3, averaging 3.04 bits per weight |
| Artifact size | 273 GiB |
A Mixture-of-Experts model stores many specialized feed-forward networks but activates only a subset for each token. GLM-5.3 selects 8 of its 256 routed experts per token and also uses a shared expert, reducing active computation even though every expert must remain available in memory.
The quantizer assigns more bits to components considered sensitive to compression and fewer bits to the routed experts that account for much of the model’s size.
| Component | Precision |
|---|---|
| Attention layers | 5 bpw |
| Shared experts | 5 bpw |
| Dense MLPs | 4 bpw |
| Routed experts | 3 bpw |
| Language-model head |
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.