Goodfire Cuts AI Jailbreaks From 66 to Zero on Kimi K3

Goodfire's probe-plus-judge cascade watches Kimi K3 and GLM 5.3 activations in real time, matching LLM judges at 50x lower cost while blocking universal jailbreaks.

·
·
·
Read6 min
TypeNews
TopicSecurity · Infra
  • Goodfire shipped probe-based cybersecurity monitors for Kimi K3 and GLM 5.3 on production inference
  • Cascade reads internal activations, escalating to an LLM judge only when a threshold trips
  • Matches LLM judge recall at roughly 50x lower cost and 50x lower added latency
  • FAR.AI red team: 66 universal jailbreaks dropped to 0 across 140 attack strategies
  • 93% recall on harmful sessions at a 5.5% benign-interruption rate across 60,000 turns
  • Built on SGLang with custom kernels; throughput hit under 2% versus default serving

Goodfire has released cybersecurity monitors for Kimi K3 and GLM 5.3 that inspect each model’s internal activations while it generates tokens. The company reports harmful-session recall comparable to an LLM judge at roughly one-fiftieth the cost and latency. In external testing by FAR.AI, the monitor reduced the reported count of universal jailbreaks from 66 to zero. The approach targets a growing problem for autonomous agents: monitoring trajectories that can span millions of tokens before those agents invoke tools or modify systems.

Why every-turn judges break

Benign security work and malicious exploitation often involve the same commands, code, and tools. A useful monitor must therefore track intent across an interaction, rather than classify isolated prompts or outputs. That becomes expensive when an agent runs for thousands of turns.

Goodfire estimates that using another model to judge one million Kimi K3 turns would cost about $30,000 with Kimi K3 or roughly $200,000 with Claude Opus 5. Many production systems respond by evaluating logs asynchronously, leaving harmful tool calls unblocked while review is pending.

A cheap gate before the judge

Goodfire’s two-stage design, documented in its research post, places a lightweight activation probe ahead of an LLM judge. The probe scores internal activation vectors, the numerical representations produced inside the model during a forward pass. When its score crosses a calibrated threshold, the system escalates the interaction to the judge, which decides whether generation should continue.

The probe reuses activations already computed during inference, avoiding a second model call on most turns. Goodfire designed it to detect two signal types:

  • Localized patterns associated with a specific passage in one turn.
  • Distributed patterns that become visible only across a full turn or trajectory.

On a held-out calibration set of chat and agentic interactions, the probe recorded a 0% false-negative rate and a 95% true-negative rate. In that set, every harmful example triggered escalation, while 95% of benign examples avoided the heavier judge. Those calibration figures describe a controlled held-out set; broader evaluations produced lower recall.

Putting probes in the serving path

Live activation monitoring required changes to the inference stack. Goodfire extended SGLang to calculate probe scores inside the server, added custom kernels, and carried the scores through its radix tree, speculative decoding, and split prefill-decode pipeline. Those components handle prompt caching, draft-token generation, and the separation of prompt processing from token generation. Each worker serving a Kimi K3 replica runs the probes locally and streams monitoring events with the response.

Performance figures reported by Goodfire
Metric Judge on every turn Probe and judge cascade
Added latency without reasoning 160 ms 2 ms
Added latency with reasoning 2,750 ms 42 ms
Cost per million turns About $30,000 with Kimi K3; about $200,000 with Claude Opus 5 $65 to $185
Share of compute budget Up to 20% Under 1%
Throughput reduction Significant Under 2%

Goodfire reports that the cascade reduced judge costs by about 49 times without reasoning and 29 times with reasoning. Because scoring runs synchronously, the serving layer can stop a risky trajectory before the agent calls a tool.

Training for shifting intent

Goodfire generated training data from simulated users assigned harmful and benign goals in cyber and non-cyber environments. Some offensive-security rollouts used Prime Intellect’s verifiers, sandboxes, and tunnels, allowing code to execute without reaching the open internet.

A written policy distinguished bounded activities, such as reproducing a vulnerability in an authorized environment, from actions such as active exploitation. Simulated users could also escalate their intent during a conversation, giving the probe examples of trajectories that begin harmlessly and become dangerous.

The project builds on Goodfire’s earlier probe study. That work found reward hacking in 50% to 96% of rollouts across three open-weight models and three agentic benchmarks. Small classifiers trained on internal activations detected the behavior during inference with substantially less computation than an LLM judge.

Recall rises, false alarms remain

Goodfire evaluated the cascade on 2,400 chat and agentic sessions containing more than 60,000 turns. The test data came from Anthropic’s Petri red-teaming library, ExploitBench, a cybersecurity subset of WildChat, and benign software-engineering rollouts.

At an operating point that interrupted 5.5% of benign sessions, the cascade detected about 93% of harmful sessions and flagged none of the non-cyber traffic. On Kimi K3, harmful-session blocking rose from about 29% to 96%, while harmless interruptions increased from 5% to 9%. On GLM 5.3, harmful-session blocking rose from 36% to 91%.

  • Against Jev: The cascade reached 93% recall with $65 in judge compute per million turns and 2 ms of latency. Jev, a lower-cost decision-model judge, reached 86% recall at $600 and 200 ms.
  • Benign framing: Jev detected 12% of harmful prompts presented as benign requests.
  • Probe architecture: Standard linear probes scored about five percentage points below Goodfire’s selected design. Nonlinear probes lost up to 30 percentage points of recall on out-of-distribution data.
  • Model transfer: A monitor developed with Kimi data produced similar results on GLM 5.3 without GLM-specific training data.

External red team: 66 to zero

Goodfire gave the deployed system to FAR.AI for two days of external testing. FAR.AI began with 1,500 static jailbreak strategies and selected 140 that succeeded against the unmonitored base model. The resulting portfolio was non-adaptive, meaning the attacks were fixed in advance and did not change in response to the monitor.

Across 1,120 test interactions, the 140 strategies produced 66 universal jailbreaks and 700 successful jailbroken interactions against the raw model, according to Goodfire’s report. With monitoring enabled, the same test produced zero universal jailbreaks and 18 successful interactions.

Universal-jailbreak success ranged from 12% to 53% without monitoring, depending on the model’s reasoning effort. Per-prompt success fell from a range of 9% to 40% to a range of 0.1% to 1.6%. The remaining 18 successful interactions show that the monitor reduced attack success under this test protocol without eliminating it.

What deployment demands

Activation probes move enforcement into the inference server, where they can inspect generation continuously and intervene before a tool call. That requires access to model weights, intermediate activations, and the serving stack. Teams using a closed hosted API cannot add this form of monitoring at the application layer.

Open-model providers can integrate the probe alongside batching, caching, speculative decoding, and distributed inference. Goodfire has deployed the system with Baseten, indicating that the design can operate in a managed serving environment as well as a research setup.

Several constraints shape production use:

  • A 93% harmful-session recall rate leaves some attacks undetected.
  • The policy used for training defines acceptable security work and may require adjustment for each organization.
  • The FAR.AI evaluation used fixed attacks, so adaptive attackers who iterate against the monitor remain a separate test case.
  • False positives can interrupt benign work, making threshold selection a product and policy decision.

For teams running open-weight code-auditing assistants, DevOps agents, or autonomous penetration-testing tools, the reported results provide a concrete architecture for synchronous monitoring: score activations on every turn, escalate uncertain cases to a stronger judge, and block risky tool calls before execution.

Trending
  • No trending articles

Comments

avatar

Next Reads