DASLab Squeezes a 176.9B Coding Model Down to 32GB Hardware
ISTA's DASLab shrinks a 176.9B-parameter Qwen mixture-of-experts model to 29.6GB resident memory by pruning half its experts and quantizing the rest to 3.5 bits.
- ISTA-DASLab released Qwen3.8-Flash-Next GSQ-RCO Coder, a 176.9B MoE compressed to 58.4GB total, 29.6GB resident.
- Half the 512 experts per layer removed via RCO; remaining weights quantized to 3.5 bits per weight.
- Retains 91.3% of BF16 SWE-bench Verified score and 98.7% of LiveCodeBench v6.
- Calibration mixture targets code, agentic tool use, vision, and spatial reasoning; general capability degrades.
- Runs unmodified in llama.cpp, Ollama, LM Studio, vLLM; Apache-2.0 license inherited from base.
- Reports of 44 tok/s output on 32GB consumer hardware; based on GSQ and RCO papers.
A 176.9B coding MoE targets 32GB hardware
IST Austria’s Deep Algorithms and Systems Lab, or DASLab, has released the Qwen3.8 coder model, a coding-focused compression of Qwen’s Flash-Next mixture-of-experts model. The lab removed half of its routed experts, quantized the remaining weights to 3.5 bits per weight, and packaged part of the model as disk-backed lookup data.
Flash-Next contains 176.9 billion total parameters, but its sparse architecture activates only a small subset for each token. Sparse activation reduces computation while leaving every expert to be stored somewhere. DASLab’s release targets that storage cost, reducing the 354GB BF16 model to a 58.4GB two-shard GGUF package and reporting 29.6GB of model data that must remain memory-resident.
- Original model: 176.9B parameters and 354GB in BF16, a 16-bit weight format
- Expert count: 512 routed experts per layer, reduced to 256
- Active experts: 10 per token
- Quantization: 3.5 bits per surviving weight
- GGUF package: 58.4GB across two shards
- Reported resident data: 29.6GB, with an n-gram lookup shard available from disk
- Reported effective precision: 1.89 bits per original transformer parameter
- License: Apache 2.0, inherited from the base model
The 1.89-bit figure averages weight storage over the original transformer parameter count, including experts removed during pruning. It should not be used to calculate the complete GGUF package size, which also includes metadata and the n-gram lookup data.
Half the experts leave the router
Across 48 layers, Flash-Next contains 24,576 routed experts and selects 10 for each token. DASLab says expert matrices account for 95% of the weights considered during its search, making them the main target for reducing storage.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.