MiaAI's Qwen3.8-27B Installer Boosts JSON and Code Speed by 78%

A one-click Qwen3.8-27B serving kit for consumer NVIDIA cards gets a speedup, with code generation jumping from 180 to 275 tok/s on an RTX 5090.

·
·
·
MiaAI's Qwen3.8-27B Installer Boosts JSON and Code Speed by 78%PRO
  • MiaAI-Lab's one-click Qwen3.8-27B installer upgrades ExLlamaV3 to 1.6.0 with a DFlash2 MTP draft model.
  • RTX 5090 code generation jumps from 180 to 275 tok/s, structured output from 206 to 366 tok/s.
  • Prose generation sees a smaller bump, 103 to 113 tok/s, since prose is harder to speculatively predict.
  • Runs on 12 GB+ NVIDIA cards (Turing and newer); 2.0 bpw EXL3 quant fits 16 GB VRAM.
  • No CUDA Toolkit, Build Tools, or Git required; serves an OpenAI-compatible endpoint plus chat UI.
  • In-flight PR adds Anthropic Messages API shim and two-GPU memory splitting, but no true tensor parallelism.

Qwen3.8-27B installer boosts code and JSON throughput

Running a 27-billion-parameter model on one consumer GPU often requires matching CUDA dependencies, compiling extensions, and choosing a quantization that fits available VRAM. The MiaAI-Lab installer packages those steps for Qwen3.8-27B. Its latest update pins ExLlamaV3 1.6.0 and adds a DFlash2 draft model for speculative decoding, with the largest reported RTX 5090 gains appearing in code and structured output.

Two changes under the hood

The repository serves Qwen/Qwen3.8-27B using turboderp’s EXL3 quants, a compressed weight format designed for ExLlamaV3. It detects the installed NVIDIA card, selects a quant that fits, creates an isolated Python environment, downloads the weights, starts an OpenAI-compatible endpoint, and opens a chat interface.

Speculative decoding uses a smaller model to draft several likely tokens before the 27B model verifies them together. The previous implementation used a Multi-Token Prediction head, a small auxiliary network attached to the main model. The new DFlash2 draft model produces candidates faster, while ExLlamaV3 1.6.0 handles generation and verification.

Where the gains show up

Project-reported RTX 5090 output throughput
Workload Previous

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads