Mia's AI Lab Ships GLM-5.3-Flash With Uncensored Weights and 16x Faster Responses
Mia's AI Lab shipped an abliterated build of GLM-5.3-Flash tuned for DGX Spark, trading roughly 6.5% quality for near-total refusal bypass.
- Mia's AI Lab released an abliterated GLM-5.3-Flash EXL3 4bpw quant for DGX Spark.
- Refusal removal via drowzeys' o_proj tensors on layers 15-43 plus MTP layer 45.
- Reported ~6.5% quality loss; 32/32 Refusal32 bypass measured on the NVFP4 parent.
- Toggle between standard and abliterated weights with a single ABLIT=1 flag.
- Next-turn chat latency cut from ~16s to ~1s; parallel-agent memory corruption fixed.
- Gated download requiring agreement to responsible-use terms; MIT license inherited from Z.AI.
Mia’s GLM-5.3-Flash branch adds selectable refusal-modified weights
Mia’s AI Lab has released an abliterated variant of its GLM-5.3-Flash quantization for two NVIDIA DGX Spark systems. The checkpoint preserves the parent model’s layout, allowing existing deployments to select the modified weights with one configuration flag. Updated serving code also reduces next-turn latency, prevents shared-cache corruption, and fixes the non-reasoning chat template.
Abliteration is a weight-editing technique that suppresses model refusals without a conventional fine-tuning run. Researchers estimate an activation direction associated with refusal behavior, then modify selected weight matrices to weaken that direction. The method can reduce refusals while introducing capability loss or unexpected behavior, so its effects require task-specific evaluation.
The edit stays inside attention output projections
The abliterated checkpoint contains an 88-billion-parameter mixture-of-experts model quantized to an average of four bits per weight in EXL3 format. Mixture-of-experts models route each token through a subset of expert weights, while EXL3 reduces memory use for ExLlama-based inference.
| Model section | Tensor source | Change |
|---|---|---|
| Transformer layers 15 to 43 | Keys tensors | self_attn.o_proj replaced |
| Multi-token-prediction layer 45 | Keys tensors | self_attn.o_proj replaced |
| Layers 0 to 14 and layer 44 | Parent checkpoint | Unchanged |
| Experts, vision layers, QKV projections, embeddings, and output head | Parent checkpoint | Unchanged |
The replacement source tensors come from drowzeys and were derived from dealignai’s uncensored NVFP4 build. Drowzeys reports bypassing all 32 prompts in the Refusal32 test, with zero refusals and zero garbled responses. Those results were measured on the NVFP4 parent and do not establish equivalent behavior for the EXL3 conversion.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.