SC117 Strips Qwen3.8-Flash-Next Refusals Without Touching 1,079 Tensors

A refusal-removed build of Qwen3.8-Flash-Next transplants 144 tensors into an already-quantized GGUF without recomputing a single learned value.

·
·
·
SC117 Strips Qwen3.8-Flash-Next Refusals Without Touching 1,079 TensorsPRO
Read2 min
TypeModel
TopicGpus · Llms
  • SC117 released an abliterated GGUF of Qwen3.8-Flash-Next with 900K+ downloads.
  • 144 refusal-direction tensors across 48 layers transplanted without recomputing GSQ's learned quantized values.
  • Four tiers from Q2_0 (38 GB) to IQ3_S (55 GB), only 0.25-0.44 GB larger than upstream.
  • Decode speed 104.4 tok/s vs upstream 104.5, 262K context, runs on 12 GB GPU via Strata.
  • Works unmodified in llama.cpp, Ollama, LM Studio, vLLM; also supports vision and speculative decoding.
  • Refusal-removed with no built-in guardrails; deploy only behind your own moderation layer.

A 144-tensor refusal edit avoids full requantization

SC117 has published an abliterated Qwen3.8-Flash-Next GGUF designed to reduce learned refusal behavior while preserving the upstream model’s quantized values outside the edited tensor set. The release changes 144 projection tensors across 48 layers, retains 1,079 non-target tensors byte for byte, and re-encodes nine targets with their learned GSQ scales.

At publication, the SC117 repository had passed 900,000 downloads and was near the top of Hugging Face’s trending list. Its technical contribution is a way to apply a narrow post-training edit to a model whose learned quantization cannot be reproduced by a conventional BF16-to-GGUF conversion.

Why full requantization loses the original

Abliteration is the informal name for a low-rank weight edit that projects an activation direction associated with refusals out of selected model matrices. A rank-1 edit uses one vector direction, making the correction much smaller and more structured than a general weight update.

A conventional workflow applies the edit to BF16, a 16-bit floating-point checkpoint, and then quantizes the entire model again. That process cannot preserve the exact low-bit values in this release because the upstream quantization was learned during optimization.

The upstream ISTA-DASLab build combines GSQ and RCO. GSQ learns grid assignments and per-group scales for each tensor through a Gumbel-Softmax relaxation. RCO then selects a quantization type for each tensor while keeping the complete model within a size budget. The model uses a mixture-of-experts architecture, which routes tokens through selected expert feed-forward blocks instead of activating every expert for every token.

Quantized values produced by GSQ-RCO depend on the learned codes, scales, and tensor-level format choices. Running a standard quantizer over BF16 weights would produce a different low-bit checkpoint and discard those optimized parameters.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads