OrcaRouter's OrcaSAQ-2 Shrinks a 27B Security AI 71% to Fit One GPU
OrcaRouter compressed a 27B uncensored cyber-focused Qwen model from 54.7 GB to 15.7 GB while keeping 94.4% token agreement with the original.
- OrcaSAQ-2 Cyber 27B compresses Qwen3.8-27B Uncensored from 54.7 GB to 15.7 GB.
- Reports +0.80% perplexity, 94.4% Top-1 agreement, and 0.020 mean KLD versus BF16 on WikiText-2.
- Preserves the full 262K context via hybrid Gated DeltaNet plus full-attention architecture.
- Runs at 27.6 tok/s single-stream on a 24 GB GPU with DFlash2 lossless speculative decoding.
- Uncensored abliterated checkpoint aimed at defensive red teaming, vulnerability research, and local security coding.
- Apache-2.0 licensed GGUF, works with llama.cpp, Ollama, LM Studio, vLLM, and Docker Model Runner.
OrcaRouter Shrinks an Uncensored 27B Cyber Model to 15.7 GB
OrcaRouter has released OrcaSAQ-2 Cyber 27B Uncensored GGUF, a compressed Qwen3.8-27B derivative built for local security work. According to the model card, quantization reduces the checkpoint from 54.7 GB in BF16 to 15.7 GB, allowing the weights to run on a single 24 GB or 16 GB consumer GPU.
Local execution lets security teams analyze proprietary source code, live logs, suspicious binaries, and unpatched vulnerabilities without sending that material to a hosted model provider. The checkpoint also has its learned refusal behavior removed, making it suitable for authorized research that may involve exploit-adjacent prompts. Accuracy, access control, and operational safety remain the deployer’s responsibility.
Compression Holds Close on WikiText-2
OrcaRouter evaluated the quantized checkpoint against the BF16 reference on 16,376 predicted WikiText-2 tokens. The published results show a small change in general language-modeling performance alongside a 71.3% reduction in storage.
| Metric | BF16 reference | Quantized model | Reported result |
|---|---|---|---|
| Perplexity | 5.6532 | 5.6961 | +0.80% |
| Top-1 token agreement | Baseline | 94.4% | 5.6% of top choices differ |
| Mean KLD | Baseline | 0.020 | Small distribution shift |
| Checkpoint size | 54.7 GB | 15.7 GB | 71.3% smaller |
| Maximum context | 262,144 tokens | 262,144 tokens | Architecture unchanged |
Perplexity measures next-token prediction error, with lower values indicating a closer fit to the evaluation text. Top-1 agreement means the compressed model selected the same next token as BF16 about 94 times out of 100. Mean Kullback-Leibler divergence, or KLD, measures the gap between the models’ full token-probability distributions.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.