SuperQwen3.8 Runs Uncensored on One Box at 1.77x Faster Speeds

A community fine-tune of Qwen3.8-27B strips refusals, fixes the resulting weight damage, and ships in four formats tuned for local inference.

·
·
SuperQwen3.8 Runs Uncensored on One Box at 1.77x Faster SpeedsPRO
Read2 min
TypeModel
TopicGpus · Llms
  • SuperQwen3.8-27b-abliterated-GGUF released by Jun Song: uncensored, multimodal, weight-repaired Qwen3.8-27B build.
  • Rank-4 OBLITERATUS refusal-subspace edit; vision tower and MTP weights preserved exactly.
  • Q4_K_M target plus MTP draft and vision projector total 17.56 GiB, fits one DGX Spark.
  • C1 decode: 12.5 tok/s baseline, 22.1 tok/s with native MTP K=3 speculative decoding.
  • 262K context metadata, verified 46,235-token needle retrieval, Apache-2.0 license.
  • Four native formats also available: BF16, NVFP4, GGUF, and MLX 4-bit.

A new community release is making the rounds on Hugging Face: SuperQwen3.8-27b-abliterated-GGUF, a repackaged and behavior-modified variant of Qwen3.8-27B that ships in a single-box GGUF bundle designed to run on one NVIDIA DGX Spark. The author, Jun Song, describes it as an uncensored, weight-repaired, multimodal build with a 1M token context, and the release covers four native formats: BF16, NVFP4, GGUF, and MLX 4-bit.

What abliteration actually did here

The base of this release is Jiunsong/SuperQwen3.8-27b-abliterated, which applies a measured rank-4 OBLITERATUS refusal-subspace edit to Qwen/Qwen3.8-27B while preserving the official vision tower and MTP weights exactly. Abliteration is a direct weight edit that identifies the internal direction a model uses to represent refusal and subtracts it, so the model complies with prompts the parent would decline. No LoRA or inference-time adapter is required, so the modification is baked into the weights.

Naive abliteration usually damages the model. The originating announcement calls out that the super-tune takes much more time than normal abliterated variants, because fixing weights and running evals takes about a week, and Song says the damage was patched using an agent swarm before release. According to the model card, the result is a capability floor of 7/8, with tool call and vision gates passing, plus reasoning that stops on all nine deterministic tasks at default, low, medium, and xhigh.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads