OpenAI's Private GPT-6 Sol Tops Cyber Defense Rankings With Zero Refusals

A reduced-guardrail OpenAI model accessed through a private program jumps 32 points on cyber defense benchmarks while costing a fraction of rivals.

·
·
·
OpenAI's Private GPT-6 Sol Tops Cyber Defense Rankings With Zero Refusals
Read4 min
TypeNews
  • GPT-6 Sol (Daybreak Blue, max) tops the Artificial Analysis Cyber Index as a trusted-access entry.
  • Scored 32 points higher than public GPT-6 Sol (max), with zero safety blocks across all tasks.
  • Costs $1.77 per task versus $11.67 for Grok 4.7 (xhigh), the prior leader.
  • Largest gains came on CyberGym-E2E, the benchmark with the most safety refusals.
  • Index covers defensive cyber only: finding, reproducing, and patching vulnerabilities in real code.
  • Cyber Index Alliance members include Collinear AI, IBM, NVIDIA, and Vercel.

Private GPT-6 variant takes Cyber Index lead with zero refusals

Artificial Analysis has expanded its Cyber Index to include restricted-access models. The first entrant, GPT-6 Sol (Daybreak Blue, max), is available through OpenAI’s private Daybreak program and now leads the ranking for AI-assisted cyber defense.

The Daybreak Blue variant scored 32 points above the publicly available GPT-6 Sol (max). Artificial Analysis attributes almost all of that gain to safety behavior: the private model recorded zero refusals across the evaluation suite, while the public version declined tasks that resembled offensive security work.

Three tests cover the defensive loop

The Cyber Index measures defensive work across three stages: finding vulnerabilities in source code, reproducing and validating them, and producing patches that preserve existing functionality. The suite excludes requests to build working exploits.

Benchmark Origin Metric
CWE-Bench-AA Collinear AI’s CWE-bench pass@1
DeepsecBench-AA Vercel’s DeepSecBench F2, median of three runs
CyberGym-E2E-AA Berkeley Center for Responsible, Decentralized Intelligence pass@1

The overall score is the equally weighted mean of the three metrics, each normalized to a 0-to-100 scale. Pass@1 measures first-attempt success. The F2 score combines precision and recall while placing more weight on recall, which penalizes missed vulnerabilities more heavily.

Refusals explain the 32-point gap

The largest improvement appeared on CyberGym-E2E-AA, where the public model accumulated the most safety refusals. Reproducing a crash in a real codebase can resemble exploit development to a safety classifier, even when the requested output is a validated bug report and patch.

Every refusal receives a score of zero. The index therefore captures both technical capability and the provider policies that determine whether users can invoke it. A model that could complete a task but declines it receives the same task score as one that attempts the work and fails.

A 6.6-fold cost gap

Model Access Reported result Average task cost
GPT-6 Sol (Daybreak Blue, max) Private New leader; 32 points above the public variant $1.77
GPT-6 Sol (max) Public 32 points below Daybreak Blue Not stated
Grok 4.7 (xhigh) Public 56; previously top-ranked $11.67
MiMo-V2.6-Pro Public 56 Not stated
GPT-6 Luna (max) Public 53 Not stated

At the reported averages, a Grok 4.7 benchmark task costs about 6.6 times as much as a Daybreak Blue task. Repeated analysis across large repositories or many candidate vulnerabilities amplifies that difference, although production costs will also depend on prompt size, repository structure, and retry behavior.

One index, two access tiers

The expanded leaderboard now reports capability under two access policies:

  • Public access: models that developers can buy and deploy through generally available products.
  • Trusted access: models offered to vetted partners under private programs with different safety and operating controls.

The separation preserves a practical view of deployable performance while exposing the capability that safety refusals suppress. On this evaluation suite, the 32-point difference provides a benchmark-specific estimate of that constraint.

The Cyber Index Alliance includes Collinear AI, IBM, NVIDIA, and Vercel. Members contribute datasets, research, and methodological input, while Artificial Analysis runs the evaluations independently.

How to evaluate a model for production

Teams selecting a model for security tooling can use the update to tighten their own evaluations:

  1. Record refusals, attempted failures, and successful fixes separately so policy blocks remain visible in the results.
  2. Test representative repository code because refusal behavior can vary with vulnerability type, prompt wording, and the amount of exploit-adjacent context.
  3. Calculate total spend per successful result, rather than relying only on average task cost.
  4. Base capacity plans on the access tier the production system can actually use.

Teams without Daybreak access should use the public tier for procurement and include refusal rates in acceptance tests. Teams admitted to the private program can compare the capability gain with its access limits and operating requirements.

Trending
  • No trending articles

Comments

avatar

Next Reads