Blackfrost-AI's CYBER-FROST Brings a 180B Security Model to Local Machines
A 180B security-focused Mixture of Experts model is now available in eight local-friendly GGUF quants, with speeds up to 2.4 tokens per second on a single GPU.
- CYBER-FROST-3.8-GGUF ports Blackfrost-AI's 180B security-focused MoE to eight local quant sizes.
- Built on Qwen3.8-Flash-Next with 512 experts, 10 active per token, ~6B active parameters.
- Fine-tuned to reduce false refusals on authorized offensive and defensive security tasks.
- Sizes range from 80 GB (Q2_K_S) to 123 GB (UD-Q4_K_XL); MXFP4_MOE variant hits 1.5 tok/s.
- MTP speculative draft head fails to load in llama.cpp due to a tensor naming mismatch.
- 262k configured context, but only 512-token smoke tests validated; quantized KV cache crashes it.
CYBER-FROST brings a 180B security model to GGUF
CYBER-FROST-3.8-GGUF packages Blackfrost-AI’s Qwen-based cybersecurity fine-tune as local GGUF files for llama.cpp. The release offers eight quantized variants, with reported file sizes from 80 GB to 123 GB, so developers can balance model fidelity, storage, memory pressure, and decode speed on a high-memory workstation. Security teams gain a way to keep sensitive engagement data off hosted APIs, although benchmark coverage and long-context behavior remain limited.
A 180B model that activates 6B
The underlying BF16 checkpoint is a Mixture of Experts model based on Qwen3.8-Flash-Next. It contains about 180 billion parameters but routes each token through 10 of 512 experts plus a shared expert, activating roughly 6 billion parameters per token. That routing reduces computation per forward pass, while the full expert pool still requires substantial storage and memory bandwidth.
| Component | Specification |
|---|---|
| Text stack | 48 blocks with hybrid linear and full attention |
| Full-attention cadence | Every fourth block |
| Hidden size | 2,560 |
| Attention heads | 24 query heads and 2 key-value heads |
| Expert routing | 512 routed experts, 10 selected per token, plus one shared expert |
| Parameters | About 180B total and 6B active per token |
| Configured context | Up to 262,144 tokens |
Security refusals are the tuning target
Security prompts frequently combine terms associated with both legitimate testing and malicious activity. General-purpose assistants may refuse requests involving exploit validation, malware analysis, incident response, or detection engineering even when the operator is working within an approved scope. Blackfrost’s BF16 checkpoint was fine-tuned to reduce those false refusals for authorized offensive and defensive workflows.
The model card lists coverage across several security domains:
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.