Alibaba's Qwen2.5-Coder Hits 775K Downloads Running on 8GB GPUs

A community 4-bit AWQ build of Qwen2.5-Coder-7B-Instruct is pulling in three quarters of a million downloads, shrinking the model to under 6GB of VRAM.

·
·
·
Alibaba's Qwen2.5-Coder Hits 775K Downloads Running on 8GB GPUsPRO
  • Community AWQ 4-bit quant of Qwen2.5-Coder-7B-Instruct has crossed 775K downloads on Hugging Face.
  • Runs in ~5.7GB VRAM with a 32K context, usable on 8GB consumer GPUs.
  • Base model trained on 5.5T code tokens, covers 92 programming languages, Apache 2.0 licensed.
  • Benchmark highlights: HumanEval 87.8%, LiveCodeBench 28.6%, SWE-Bench Verified 20.3%.
  • Drop-in with Transformers and vLLM via standard model_id loading, no API fees.
  • Independent trackers rate it below frontier intelligence but strong value for local coding workflows.

A 4-bit Qwen coder passes 775,000 downloads

A community AWQ checkpoint of Alibaba’s Qwen2.5-Coder-7B-Instruct has passed 775,000 downloads on Hugging Face. The package converts the model’s weights to 4-bit precision, bringing the reported model load to about 5.7GB and enabling short-context inference on many 8GB GPUs.

The checkpoint retains the architecture, tokenizer, instruction tuning, and Apache 2.0 license of the source model. Its appeal comes from deployment: BF16 weights for a 7.61-billion-parameter model consume roughly 15.2GB before runtime overhead, while this build leaves some room for activations and the key-value cache on smaller cards. Hugging Face’s counter includes repeated and automated file pulls, so the total cannot establish a unique-user count.

Inside the AWQ checkpoint

AWQ, short for Activation-aware Weight Quantization, uses calibration activations to identify sensitive weight channels and select scaling factors before converting weights to lower precision. The method reduces quantization error while keeping the stored weights compact. The 4-bit label applies to weights; activations and the key-value cache generally retain higher precision unless the inference engine configures them separately.

Specification Details
Parameters 7.61 billion
Transformer layers 28
Attention 28 query heads and 4 shared key-value heads
Architecture Qwen2 with RoPE, SwiGLU, RMSNorm, and grouped-query attention
Quantization 4-bit AWQ
Weight download About 5.3GB across four files
Reported model VRAM About 5.7GB before workload-dependent overhead
Configured context 32,768 tokens
Extended context Up to 131,072 tokens using the parent model’s YaRN configuration
License Apache 2.0, including commercial use subject to its terms

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads