Alibaba's Qwen2.5-Coder Hits 775K Downloads Running on 8GB GPUs
A community 4-bit AWQ build of Qwen2.5-Coder-7B-Instruct is pulling in three quarters of a million downloads, shrinking the model to under 6GB of VRAM.
- Community AWQ 4-bit quant of Qwen2.5-Coder-7B-Instruct has crossed 775K downloads on Hugging Face.
- Runs in ~5.7GB VRAM with a 32K context, usable on 8GB consumer GPUs.
- Base model trained on 5.5T code tokens, covers 92 programming languages, Apache 2.0 licensed.
- Benchmark highlights: HumanEval 87.8%, LiveCodeBench 28.6%, SWE-Bench Verified 20.3%.
- Drop-in with Transformers and vLLM via standard model_id loading, no API fees.
- Independent trackers rate it below frontier intelligence but strong value for local coding workflows.
A 4-bit Qwen coder passes 775,000 downloads
A community AWQ checkpoint of Alibaba’s Qwen2.5-Coder-7B-Instruct has passed 775,000 downloads on Hugging Face. The package converts the model’s weights to 4-bit precision, bringing the reported model load to about 5.7GB and enabling short-context inference on many 8GB GPUs.
The checkpoint retains the architecture, tokenizer, instruction tuning, and Apache 2.0 license of the source model. Its appeal comes from deployment: BF16 weights for a 7.61-billion-parameter model consume roughly 15.2GB before runtime overhead, while this build leaves some room for activations and the key-value cache on smaller cards. Hugging Face’s counter includes repeated and automated file pulls, so the total cannot establish a unique-user count.
Inside the AWQ checkpoint
AWQ, short for Activation-aware Weight Quantization, uses calibration activations to identify sensitive weight channels and select scaling factors before converting weights to lower precision. The method reduces quantization error while keeping the stored weights compact. The 4-bit label applies to weights; activations and the key-value cache generally retain higher precision unless the inference engine configures them separately.
| Specification | Details |
|---|---|
| Parameters | 7.61 billion |
| Transformer layers | 28 |
| Attention | 28 query heads and 4 shared key-value heads |
| Architecture | Qwen2 with RoPE, SwiGLU, RMSNorm, and grouped-query attention |
| Quantization | 4-bit AWQ |
| Weight download | About 5.3GB across four files |
| Reported model VRAM | About 5.7GB before workload-dependent overhead |
| Configured context | 32,768 tokens |
| Extended context | Up to 131,072 tokens using the parent model’s YaRN configuration |
| License | Apache 2.0, including commercial use subject to its terms |
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.