GLM-4.7-Flash Roleplay Fine-Tune Lands Calibrated GGUF Builds for Local Inference
A community imatrix-quantized GGUF build of an uncensored GLM-4.7-Flash fine-tune brings local Chinese and English roleplay to consumer hardware.
- Community imatrix GGUF quant of a NSFW fine-tune of Z.ai's GLM-4.7-Flash.
- Base model is a 30B-A3B MoE with ~3.6B active parameters and 128K-200K context.
- Runs on 24GB VRAM at Q4_K_M via llama.cpp, LM Studio, Ollama, Jan.
- MIT licensed, supports Chinese and English, marked not-for-all-audiences.
- Requires post-January llama.cpp build to avoid a scoring_func bug causing looping.
- Target use is uncensored roleplay and creative writing on consumer hardware.
GLM-4.7-Flash roleplay fine-tune gets calibrated GGUF builds
A community upload on Hugging Face now packages a roleplay-focused GLM-4.7-Flash derivative as imatrix-calibrated GGUF files. The GGUF repository targets llama.cpp and compatible applications, giving local-inference developers ready-to-load low-bit builds without requiring a separate LoRA merge or conversion step.
The release matters primarily for developers evaluating Chinese character chat, reduced-refusal behavior, or quantization quality on sparse mixture-of-experts models. Its NSFW label describes the uploader’s intended use, while actual behavior, language quality, and safety characteristics still require application-specific testing.
| Component | Details |
|---|---|
| Base | Z.ai GLM-4.7-Flash, a 30B-A3B mixture-of-experts model |
| Derivative | Chinese-centered character roleplay fine-tune with an NSFW designation |
| Format | GGUF files calibrated with an importance matrix |
| Primary runtime | llama.cpp, plus applications built on compatible GGUF loaders |
| Context | Up to roughly 200K tokens in the upstream model, subject to runtime memory and workload quality |
| License metadata | MIT; downstream users should also review the upstream model and fine-tune terms |
Calibration targets the low-bit tradeoff
Repository metadata traces the derivative to character-roleplay work from the author of the linked
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.