DavidAU's Qwen3.8 Fine-Tune Beats the Base Model at a Tenth of Thinking Tokens

A community fine-tune of Qwen3.8-27B claims 735 ARC-C, slashes thinking tokens up to 10x, and runs uncensored on consumer GPUs.

·
·
DavidAU's Qwen3.8 Fine-Tune Beats the Base Model at a Tenth of Thinking TokensPRO
  • DavidAU released a multi-stage fine-tune of Qwen3.8-27B claiming 735 ARC-C and 882 ARC-E scores.
  • TURBO branding reflects thinking token reductions of one-half to one-tenth versus base Qwen3.8.
  • Built with COLD FUSION (custom GAIN method plus Unsloth) and Fable Fusion 711 training techniques.
  • De-censored via Heretic with Arbitrary-Rank Ablation, dropping refusals from 99/100 to 11/100.
  • 4-bit quant reportedly hits 99% of 8-bit performance; MTP variants exceed 90 t/s on a 5090.
  • Apache 2.0, 256k context, vision-capable, runs in llama.cpp, Ollama, LM Studio, and vLLM.

A community fine-tune of Qwen3.8-27B is climbing the Hugging Face trending charts with aggressive claims: higher benchmark scores than the base model, a fraction of the reasoning tokens, and a footprint that fits on a single consumer GPU. Qwen3.8-27B-TURBO-Fable-Cold-Fusion from DavidAU is a multi-stage tuned, de-censored, GGUF-packaged variant built outside any major lab.

What the benchmarks actually say

The headline pitch is a jump on ARC-Challenge and ARC-Easy. At 8-bit (mxfp8), the model reports 0.735 ARC-C and 0.882 ARC-E, versus the base Qwen3.8-27B at 0.591 and 0.782. At 4-bit (mxfp4) those numbers land at 0.719 ARC-C and 0.887 ARC-E, close enough to 8-bit performance that the author frames 4-bit as roughly 99% of 8-bit quality. For anyone running locally on 16-24 GB of VRAM, that gap matters. The benchmarks were run by the author's collaborator Nightmedia using standard multiple-choice suites (ARC, BoolQ, HellaSwag, OBQA, PIQA, WinoGrande), not a third party, and cover classification-style tasks rather than modern agentic or coding evaluations.

Slashing the thinking tokens

Qwen3.8 defaults to heavy chain-of-thought reasoning, and this fine-tune directly targets that overhead. The TURBO designation covers reductions in thinking tokens ranging from one-half to as low as one-tenth of the original, across all three reasoning modes (xhigh, medium, low), while claiming to preserve output detail and quality. Multi-turn refinement stages reportedly see the steepest cuts, sometimes reaching one-fifth the size of typical Qwen thinking blocks.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads