Goodfire's Silico Lets an Agent Crack Open AI Black Boxes Autonomously

Goodfire's Silico goes public: an AI agent that autonomously runs interpretability experiments at 2.8 trillion parameter scale, for $1,000/month

·
·
  • Public launch: Goodfire's Silico platform is now publicly available at $1,000/month for individual researchers, with 50% early-signup discounts and grants for AI safety and life sciences researchers.
  • Autonomous research agent: Silico plans and runs long-horizon interpretability and training experiments autonomously across GPU clusters, including at 2.8 trillion parameter scale (Kimi K3).
  • Hallucination reduction benchmark: Silico reproduced Goodfire's RLFR method in 2 days, cutting hallucinations in Qwen3-8B by 37%; Goodfire's own RLFR achieves up to 58% reduction at ~90x lower cost than LLM-as-judge.
  • Well-funded unicorn: Goodfire has raised $209M total, including a $150M Series B at a $1.25B valuation, backed by B Capital, Menlo Ventures, Anthropic, Eric Schmidt, and Salesforce Ventures.
  • Open-weights only: Silico requires model weight access and cannot inspect closed models like GPT-4 or Gemini.
  • Early adopters: Arc Institute, Mayo Clinic, Microsoft, Rakuten, and Prime Intellect are already using Silico for research across LLMs, genomics, and drug discovery.

Silico, Goodfire's platform for autonomous AI research, is now publicly available. The pitch is straightforward but ambitious: instead of babysitting GPU jobs and manually running interpretability experiments, you hand a research goal to an agent and it plans, executes, and returns results you can build on. That includes training runs, model probing, paper replication, and failure diagnosis, all without a human in the loop for each step.

What mechanistic interpretability actually means

Mechanistic interpretability is the practice of reverse-engineering a neural network to figure out what it has learned and why it behaves the way it does, by mapping the internal features and circuits that drive its outputs. Think of it as neuroscience for AI models. Goodfire sits alongside Anthropic, OpenAI, and Google DeepMind as one of a small handful of organizations pioneering this technique. MIT Technology Review named it one of its 10 Breakthrough Technologies of 2026.

Until now, this kind of work required a team of specialized researchers. Silico is Goodfire's bet that an agent can do most of that work autonomously. According to CEO Eric Ho, the key enabler was the maturation of AI agents themselves: "Agents are now strong enough to do a lot of the interpretability work that we were doing using humans."

What Silico actually does

Goodfire claims Silico is the first off-the-shelf tool of its kind that can help developers debug every stage of the development process, from dataset construction to model training. The platform bundles four core capabilities:

  • Model exploration: Visualize architecture, train sparse autoencoders (SAEs) and probes, map neural geometry, and test causal hypotheses about what your model has learned.
  • Failure diagnosis: Trace regressions and unexpected behavior to undertraining, information bottlenecks, feature collapse, spurious correlations, or dataset artifacts.
  • Training: Run SFT, DPO, and RL experiments, compare checkpoints, and test targeted interventions.
  • Paper replication: Give Silico a paper to reproduce or extend. It plans and runs the experiments, then compares results against the original.

Silico's core is an agent that autonomously plans and runs interpretability experiments on model internals. It also handles infrastructure: coordinating experiments across GPU clusters, monitoring every run, and keeping long-running work moving without constant supervision. Goodfire specifically notes this lets Silico operate at Kimi K3's 2.8 trillion parameter scale.

The benchmark numbers

Goodfire ran Silico against their own research to demonstrate its capabilities. The flagship example is RLFR (Reinforcement Learning from Feature Rewards), a technique that uses lightweight probes on a model's internal representations as reward signals for reinforcement learning, applied to hallucination reduction. Rather than paying for an expensive LLM-as-judge to flag bad outputs, you train a small probe that reads the model's internals directly and feeds that signal back into RL training.

Goodfire's team spent months developing RLFR. Silico reproduced it in two days, reducing hallucinations in Qwen3-8B by 37% without capability loss. The underlying method itself cuts hallucinations by 58%, at roughly 90x lower cost per intervention than LLM-as-judge, with no degradation on standard benchmarks.

The replication track record extends further. Silico autonomously reproduced months of research across three projects in hours to days:

  • J-space context extension
  • RLFR hallucination reduction
  • PICASSO cancer prediction interpretation

Applying sparse-feature probing to protein language models, Silico also surfaced internal subspaces whose activations track known protein structures, without any supervision pointing it there.

The team and the money

Goodfire is an AI interpretability research lab and public benefit corporation founded by Eric Ho (CEO), Dan Balsam (CTO), and Tom McGrath (Chief Scientist). McGrath previously worked on Google DeepMind's interpretability team. The company has raised $209M across three rounds: a $7M seed, a $50M Series A led by Menlo Ventures with Anthropic participating, and a $150M Series B led by B Capital at a $1.25 billion valuation. The Series B included DFJ Growth, Salesforce Ventures, and Eric Schmidt, among others.

The early customer list is already notable. Teams at Arc Institute, Mayo Clinic, Microsoft, Rakuten, and Prime Intellect are listed as users on the Silico product page. Vincent Weisser, CEO of Prime Intellect, described Silico as the most useful agent he has used for autonomous AI research.

Pricing and access

Silico is priced at $1,000 per month for individual researchers, with full access to interpretability and training capabilities. Goodfire is onboarding researchers in waves. Two paths exist for teams who cannot afford that rate: a 50% early-signup discount, and grants specifically for researchers in AI safety and life sciences. Goodfire also noted that long-horizon experiments are expensive by nature and that they are actively working to bring costs down.

One hard constraint worth flagging: Silico requires access to model weights and cannot inspect closed models like GPT or Gemini. It is built for teams who own or have access to open-weight models.

Beyond debugging: surgery on model behavior

"We saw this widening gap between how well models were understood and just how widely they were being deployed," says Ho. Silico is designed to close that gap by giving smaller firms and research teams access to interpretability tools that have historically lived inside a handful of frontier labs.

The life sciences work is worth watching closely. Goodfire recently identified a novel class of Alzheimer's biomarkers by applying interpretability techniques to an epigenetic model, described as the first major scientific finding obtained by reverse-engineering a foundation model. Goodfire also decoded the internal representations of Arc Institute's Evo 2 genomic model, finding features that map onto biological concepts from coding sequences to protein secondary structure, with results published in Nature.

The deeper shift Silico represents is from AI training as guesswork to AI training as targeted engineering. Using the platform, developers can adjust parameters connected to individual neurons to boost or suppress specific behaviors. In one example, Goodfire researchers asked a model whether a company should disclose that its AI behaves deceptively in 0.3% of cases. The model said no, citing negative business impact. By examining the model's internals, researchers found that boosting neurons associated with transparency flipped the answer from no to yes nine out of ten times. That kind of surgical control has simply not been accessible outside frontier labs before now.

Whether Silico's results hold up at scale and across diverse workloads beyond Goodfire's own demos remains to be seen. With $209M in funding, a Nature publication, and early adopters ranging from Mayo Clinic to Microsoft, Goodfire is making a credible case that the era of training models by feel is ending.

Trending
  • No trending articles

Comments

avatar

Next Reads