Goodfire Builds Sequence-Aware Biosecurity Monitors to Guard AI Biology Agents

Goodfire's protein-embedding monitors catch more harmful biology requests with fewer false refusals, staying 3-5x more robust against adversarial attacks.

·
·
Goodfire Builds Sequence-Aware Biosecurity Monitors to Guard AI Biology Agents
  • Goodfire released biosecurity monitors using protein-model embeddings to screen agent requests involving biological sequences.
  • Outperforms frontier model safeguards: catches more harmful requests with fewer false refusals on dual-use tasks.
  • Reported 3-5x more robust to adversarial attacks like paraphrasing and fragmentation than standard screening.
  • Runs in milliseconds per sequence, comparable to database lookup, enabling real-time deployment in workflows.
  • Reached ~0.98 AUROC on paraphrased toxins when parents were in the screening database.
  • Available via partnerships with frontier labs and inference providers, not a public API.

Goodfire builds sequence-aware monitors for biology agents

Goodfire has published a technical report on biosecurity monitors designed to sit between AI agents and biological tools. The system extracts protein sequences from requests and classifies them before a workflow redesigns, fragments, or forwards them. In Goodfire’s benchmarks, this embedding-based approach caught more concerning inputs while rejecting fewer benign ones than prompt-level safeguards and conventional sequence screening, with inference measured in milliseconds.

That tradeoff matters as AI agents gain access to protein-design software, lab automation, and synthesis ordering systems. Requests such as “refold this protein” or “split this sequence into synthesis fragments” describe routine research tasks, yet their safety depends heavily on the supplied sequence. General-purpose language models often treat those sequences as opaque strings and rely on the surrounding instructions instead.

Benign instructions can carry risky sequences

Prompt-level safeguards tend to classify the requested operation, which creates high refusal rates for legitimate work or gaps that adversaries can exploit with ordinary-looking instructions. Inverse folding, which generates a sequence for a target three-dimensional structure, can also produce redesigned proteins that preserve predicted structure while changing many residues. Synthesis workflows introduce another challenge by dividing long sequences into shorter fragments.

Goodfire’s monitor converts the extracted sequence into an embedding, a numerical representation learned by a protein language model. These representations capture statistical relationships associated with protein structure and function. A downstream classifier uses that information to assign a risk score, allowing it to detect relationships that exact string matching and profile hidden Markov model lookups may overlook.

  1. The monitor extracts a biological sequence from the request.
  2. A protein language model converts the sequence into an embedding.
  3. A classifier scores the embedding for similarity to concerning proteins.
  4. The hosting platform applies its chosen policy, such as allowing, reviewing, or blocking the request.
Sequence-aware monitoring between AI agents and biological research tools
Goodfire’s proposed monitor adds sequence-level screening between AI agents and downstream biological tools.

A controlled test isolates the sequence

Goodfire built a dual-use benchmark in which the task instructions remain constant while the input changes among known protein toxins, adversarially redesigned versions, and benign proteins. The evaluated frontier models showed similar refusal rates for concerning and benign sequences, particularly after sequence obfuscation. The embedding classifier separated those categories more consistently.

Evaluation Reported result What it tests
Fragments of held-out natural proteins Outperformed sequence-based screening at every tested fragment length Recognition when only part of the original sequence is available
Adversarial paraphrases Approximately 0.98 AUROC when the source protein appeared in the screening database Detection after substantial sequence redesign
Unseen categories Generalized across held-out protein families, toxin mechanisms, and the external Exo-Tox dataset Performance beyond closely related training examples
Runtime Milliseconds per sequence Suitability for interactive agent and tool workflows

Where the evidence stops

AUROC measures how well a classifier ranks concerning examples above benign ones across possible thresholds, with 1.0 representing perfect ranking. A reported score near 0.98 indicates strong separation in this benchmark. Production accuracy will still depend on the selected threshold, the mix of incoming traffic, and the cost assigned to false positives and missed detections.

The adversarial paraphrases were assessed through computational estimates of structure and function. Wet-lab validation was outside the evaluation, so the results do not establish whether the redesigned proteins retain biological activity. That limitation narrows the claim to detection of plausible computational attacks.

Better protein models may strengthen screening

Goodfire reports that detection improved with the capability of the underlying protein model and the quality of the classifier’s training data. The same advances that make protein-design systems more capable can therefore provide richer representations for monitoring. Continued validation will remain necessary as new protein families, design methods, and attack strategies appear.

A layered deployment can combine sequence scores with signals available elsewhere in the workflow:

  • AI agents can inspect user intent, conversation context, and requested tool calls.
  • Design tools can inspect objectives, constraints, and intermediate outputs.
  • Sequence monitors can classify the biological payload passed between systems.
  • Synthesis providers can use order history, customer identity, and expert review.

Deployment currently requires direct access

Goodfire has not released a self-service public API for the monitors. The company is working directly with frontier AI labs and inference providers, making the immediate audience model hosts, protein-design companies, and lab-automation platforms that can place screening before sensitive tool calls or synthesis requests.

Teams evaluating this approach will need to define several integration details:

  • Coverage: which prompts, tool arguments, generated sequences, and synthesis payloads receive screening.
  • Policy: which scores trigger approval, expert review, rate limits, or blocking.
  • Evaluation: how the monitor performs on local benign workloads and representative adversarial examples.
  • Operations: how decisions are logged, audited, appealed, and handled during service failures.

As biological models connect to design software, lab systems, and ordering infrastructure, sequence-level screening gives operators a more precise control than broad refusals based on task wording. Goodfire’s results support that architecture, while real-world value will depend on performance against active proteins, novel threats, and production traffic.

Trending
  • No trending articles

Comments

avatar

Next Reads