GLM-4.7-Flash Roleplay Fine-Tune Lands Calibrated GGUF Builds for Local Inference

A community imatrix-quantized GGUF build of an uncensored GLM-4.7-Flash fine-tune brings local Chinese and English roleplay to consumer hardware.

·
·
GLM-4.7-Flash Roleplay Fine-Tune Lands Calibrated GGUF Builds for Local InferencePRO
Read1 min
TypeModel
TopicLlms · Gpus
  • Community imatrix GGUF quant of a NSFW fine-tune of Z.ai's GLM-4.7-Flash.
  • Base model is a 30B-A3B MoE with ~3.6B active parameters and 128K-200K context.
  • Runs on 24GB VRAM at Q4_K_M via llama.cpp, LM Studio, Ollama, Jan.
  • MIT licensed, supports Chinese and English, marked not-for-all-audiences.
  • Requires post-January llama.cpp build to avoid a scoring_func bug causing looping.
  • Target use is uncensored roleplay and creative writing on consumer hardware.

GLM-4.7-Flash roleplay fine-tune gets calibrated GGUF builds

A community upload on Hugging Face now packages a roleplay-focused GLM-4.7-Flash derivative as imatrix-calibrated GGUF files. The GGUF repository targets llama.cpp and compatible applications, giving local-inference developers ready-to-load low-bit builds without requiring a separate LoRA merge or conversion step.

The release matters primarily for developers evaluating Chinese character chat, reduced-refusal behavior, or quantization quality on sparse mixture-of-experts models. Its NSFW label describes the uploader’s intended use, while actual behavior, language quality, and safety characteristics still require application-specific testing.

Artifact summary
Component Details
Base Z.ai GLM-4.7-Flash, a 30B-A3B mixture-of-experts model
Derivative Chinese-centered character roleplay fine-tune with an NSFW designation
Format GGUF files calibrated with an importance matrix
Primary runtime llama.cpp, plus applications built on compatible GGUF loaders
Context Up to roughly 200K tokens in the upstream model, subject to runtime memory and workload quality
License metadata MIT; downstream users should also review the upstream model and fine-tune terms

Calibration targets the low-bit tradeoff

Repository metadata traces the derivative to character-roleplay work from the author of the linked

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads