Blackfrost-AI's CYBER-FROST Brings a 180B Security Model to Local Machines

A 180B security-focused Mixture of Experts model is now available in eight local-friendly GGUF quants, with speeds up to 2.4 tokens per second on a single GPU.

·
·
Blackfrost-AI's CYBER-FROST Brings a 180B Security Model to Local MachinesPRO
  • CYBER-FROST-3.8-GGUF ports Blackfrost-AI's 180B security-focused MoE to eight local quant sizes.
  • Built on Qwen3.8-Flash-Next with 512 experts, 10 active per token, ~6B active parameters.
  • Fine-tuned to reduce false refusals on authorized offensive and defensive security tasks.
  • Sizes range from 80 GB (Q2_K_S) to 123 GB (UD-Q4_K_XL); MXFP4_MOE variant hits 1.5 tok/s.
  • MTP speculative draft head fails to load in llama.cpp due to a tensor naming mismatch.
  • 262k configured context, but only 512-token smoke tests validated; quantized KV cache crashes it.

CYBER-FROST brings a 180B security model to GGUF

CYBER-FROST-3.8-GGUF packages Blackfrost-AI’s Qwen-based cybersecurity fine-tune as local GGUF files for llama.cpp. The release offers eight quantized variants, with reported file sizes from 80 GB to 123 GB, so developers can balance model fidelity, storage, memory pressure, and decode speed on a high-memory workstation. Security teams gain a way to keep sensitive engagement data off hosted APIs, although benchmark coverage and long-context behavior remain limited.

A 180B model that activates 6B

The underlying BF16 checkpoint is a Mixture of Experts model based on Qwen3.8-Flash-Next. It contains about 180 billion parameters but routes each token through 10 of 512 experts plus a shared expert, activating roughly 6 billion parameters per token. That routing reduces computation per forward pass, while the full expert pool still requires substantial storage and memory bandwidth.

Core architecture
Component Specification
Text stack 48 blocks with hybrid linear and full attention
Full-attention cadence Every fourth block
Hidden size 2,560
Attention heads 24 query heads and 2 key-value heads
Expert routing 512 routed experts, 10 selected per token, plus one shared expert
Parameters About 180B total and 6B active per token
Configured context Up to 262,144 tokens

Security refusals are the tuning target

Security prompts frequently combine terms associated with both legitimate testing and malicious activity. General-purpose assistants may refuse requests involving exploit validation, malware analysis, incident response, or detection engineering even when the operator is working within an approved scope. Blackfrost’s BF16 checkpoint was fine-tuned to reduce those false refusals for authorized offensive and defensive workflows.

The model card lists coverage across several security domains:

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads