Goodfire's Silico Cuts Hallucinations 37% by Seeing Inside AI Models

Goodfire's Silico opens its private beta, letting an AI agent autonomously run interpretability experiments — reproducing months of research in days

·
·
AuthorGoodfire
Read2 min
SubtopicAlignment · Rlhf
  • Private beta open: Goodfire's Silico platform is now accepting requests, targeting teams actively training or fine-tuning AI models.
  • AI research agent: Silico's core is an agent that autonomously plans and runs interpretability experiments on model internals, returning results without human hand-holding.
  • Benchmark replications: Silico reproduced months of research autonomously — including J-space context extension, RLFR hallucination reduction (37% on Qwen3-8B), and PICASSO cancer prediction interpretation — in hours to days.
  • RLFR breakthrough: Goodfire's hallucination-reduction method cuts errors by up to 58% at ~90x lower cost than LLM-as-judge, with no benchmark degradation.
  • Open-source models only: Silico requires access to model weights; it cannot inspect closed models like GPT or Gemini.
  • Well-funded: Goodfire has raised $207M total; pricing is not public and access is by partnership request.

AI model training has always been a bit of a dark art. You pick a dataset, set some hyperparameters, run the training loop, and hope the model comes out the other side doing what you wanted. Goodfire's Silico is a direct challenge to that workflow. The platform, now open for private beta access, pairs frontier mechanistic interpretability techniques with an AI agent that can autonomously plan and run experiments on your model's internals , and the results it's already producing are hard to ignore.

From research lab to product

Goodfire is one of a small handful of companies, including Anthropic, OpenAI, and Google DeepMind, pioneering mechanistic interpretability , a technique that aims to understand what goes on inside an AI model when it carries out a task by mapping its neurons and the pathways between them. MIT Technology Review picked mechanistic interpretability as one of its 10 Breakthrough Technologies of 2026. Until now, this kind of work required a team of specialized researchers. Silico is Goodfire's bet that an agent can do most of that work for you.

With Silico, Goodfire is packaging up many of its in-house interpretability techniques and shipping them as a product. The tool uses agents to automate much of the complex work. "Agents are now strong enough to do a lot of the interpretability work that we were doing using humans," says CEO Eric Ho. Goodfire claims Silico is the first off-the-shelf tool of its kind that can help developers debug all stages of the development process, from building a dataset to training a model.

What the agent actually does

Silico's core is an agent that plans and runs experiments, returns results, and learns over time , all within a team workspace for training and debugging models, built on infrastructure for frontier scale. Think of it less like a dashboard and more like a junior researcher you can prompt with a high-level goal. The agent then figures out which experiments to run, executes them against your model's activations, and surfaces what it finds.

The platform's capabilities break down into a few key actions:

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves