Infatoshi Squeezes GLM-5.3's 753B Parameters Into 273 GiB for Multi-GPU Workstations

A 3-bit EXL3 quant squeezes a 753B uncensored GLM-5.3 Mixture of Experts into 273 GiB, making local inference possible on four RTX PRO 6000s.

·
·
·
Infatoshi Squeezes GLM-5.3's 753B Parameters Into 273 GiB for Multi-GPU WorkstationsPRO
  • Infatoshi released a 3.0bpw EXL3 quant of GLM-5.3-UNCENSORED, totaling 273 GiB.
  • Base is dealignai's weight-edited GLM-5.3, not a fine-tune, baked into residual-writer tensors.
  • Architecture: 753B parameters, 256 routed experts (8 active), MLA attention with DSA sparse indexer, plus MTP draft layer.
  • KL divergence vs FP8 source is 0.089; perplexity drifts from 3.302 to 3.440 on wikitext-2.
  • Runs on 4x RTX PRO 6000s (384 GB VRAM) at roughly 55-60 tok/s without speculative decoding.
  • Reports 0% refusals on HarmBench-320 at max effort, with practical 131K context ceiling on TP8 H200.

GLM-5.3’s 753B weights shrink to 273 GiB

Infatoshi has published an EXL3 release of GLM-5.3 that averages 3.04 bits per weight and occupies 273 GiB. The compression makes the 753-billion-parameter Mixture-of-Experts model practical on a high-end, multi-GPU workstation, although it remains far beyond a single consumer GPU.

The artifact quantizes dealignai’s GLM-5.3-UNCENSORED-FP8, a weight-edited version of zai-org’s original GLM-5.3. Developers evaluating it should account for both changes: the parent modifies refusal behavior, while the EXL3 conversion reduces numerical precision.

A mixed-precision squeeze

Core model and release details
Item Detail
Architecture GlmMoeDsaForCausalLM
Total parameters 753 billion
Routed experts 256, with 8 active per token
Shared experts 1
Layers 78 transformer layers and 1 MTP layer
Attention MLA with a DSA sparse indexer
Quantization EXL3, averaging 3.04 bits per weight
Artifact size 273 GiB

A Mixture-of-Experts model stores many specialized feed-forward networks but activates only a subset for each token. GLM-5.3 selects 8 of its 256 routed experts per token and also uses a shared expert, reducing active computation even though every expert must remain available in memory.

The quantizer assigns more bits to components considered sensitive to compression and fewer bits to the routed experts that account for much of the model’s size.

Precision by component
Component Precision
Attention layers 5 bpw
Shared experts 5 bpw
Dense MLPs 4 bpw
Routed experts 3 bpw
Language-model head

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads