mradermacher Brings Anime-Girl 1.5B to CPU With 20 GGUF Builds
A 1.5B parameter Qwen2-based anime chatbot lands in imatrix GGUF format, running under 1GB on plain consumer laptops.
- Imatrix GGUF quants of coderian/anime-girl-1.5B published by mradermacher on Hugging Face
- Qwen2-architecture conversational model, ~2B params, fine-tuned for anime persona chat
- Over 20 quant variants from 437 MB (IQ1_S) to 1.27 GB (Q6_K)
- Recommended Q4_K_M at 986 MB runs comfortably on CPU-only laptops
- One-line deploy via Ollama, llama.cpp, LM Studio, Jan, or Docker Model Runner
- Targeted at role-play, character chatbots, and offline hobby projects, not factual assistants
Anime-Girl 1.5B Gets Imatrix GGUF Builds for CPU Inference
Quantization maintainer mradermacher has published more than 20 imatrix-weighted GGUF versions of Anime-Girl 1.5B. The release lets developers run the conversational model through llama.cpp, Ollama, LM Studio, Jan, and other GGUF-compatible tools without preparing their own quantization pipeline or using a discrete GPU.
The upstream checkpoint uses the Qwen2 architecture and contains roughly 1.5 billion parameters. Its fine-tuning targets anime-style persona chat, placing it among the smaller models suitable for local character dialogue, prototypes, and constrained devices. The GGUF repository ranges from a 437 MB 1-bit build to a 1.27 GB 6-bit version.
How the imatrix preserves useful weights
A plain GGUF conversion can quantize weights using their numeric values alone. An importance matrix adds calibration data gathered by running sample prompts through the model and measuring which weights contribute most strongly to its activations. The quantizer uses that signal when choosing scales and rounding values, reducing error in influential parts of the network.
Calibration can improve low-bit quants, especially the IQ formats, although results depend on how closely the calibration prompts resemble the intended workload. The repository’s quality descriptions come from the publisher rather than an independent benchmark, so developers should compare persona consistency, response quality, latency, and memory use on their own prompts.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.