MiaAI's Qwen3.8-27B Installer Boosts JSON and Code Speed by 78%
A one-click Qwen3.8-27B serving kit for consumer NVIDIA cards gets a speedup, with code generation jumping from 180 to 275 tok/s on an RTX 5090.
- MiaAI-Lab's one-click Qwen3.8-27B installer upgrades ExLlamaV3 to 1.6.0 with a DFlash2 MTP draft model.
- RTX 5090 code generation jumps from 180 to 275 tok/s, structured output from 206 to 366 tok/s.
- Prose generation sees a smaller bump, 103 to 113 tok/s, since prose is harder to speculatively predict.
- Runs on 12 GB+ NVIDIA cards (Turing and newer); 2.0 bpw EXL3 quant fits 16 GB VRAM.
- No CUDA Toolkit, Build Tools, or Git required; serves an OpenAI-compatible endpoint plus chat UI.
- In-flight PR adds Anthropic Messages API shim and two-GPU memory splitting, but no true tensor parallelism.
Qwen3.8-27B installer boosts code and JSON throughput
Running a 27-billion-parameter model on one consumer GPU often requires matching CUDA dependencies, compiling extensions, and choosing a quantization that fits available VRAM. The MiaAI-Lab installer packages those steps for Qwen3.8-27B. Its latest update pins ExLlamaV3 1.6.0 and adds a DFlash2 draft model for speculative decoding, with the largest reported RTX 5090 gains appearing in code and structured output.
Two changes under the hood
The repository serves Qwen/Qwen3.8-27B using turboderp’s EXL3 quants, a compressed weight format designed for ExLlamaV3. It detects the installed NVIDIA card, selects a quant that fits, creates an isolated Python environment, downloads the weights, starts an OpenAI-compatible endpoint, and opens a chat interface.
Speculative decoding uses a smaller model to draft several likely tokens before the 27B model verifies them together. The previous implementation used a Multi-Token Prediction head, a small auxiliary network attached to the main model. The new DFlash2 draft model produces candidates faster, while ExLlamaV3 1.6.0 handles generation and verification.
Where the gains show up
| Workload | Previous |
|---|
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.