OpenAI's o3 Cracked 18 Unsolvable Rare Disease Cases in Children
OpenAI's o3 Deep Research cracked 18 previously unsolvable rare pediatric disease cases in a peer-reviewed NEJM AI study with Boston Children's Hospital
- OpenAI's o3 Deep Research and Boston Children's Hospital published a study in NEJM AI showing 18 confirmed diagnoses from 376 previously unsolved pediatric rare disease cases.
- The overall diagnostic yield was 4.8%, with the neurodevelopmental cohort hitting 10% and early psychosis reaching 13.3% -- all cases had already been through prior expert review.
- The model acted as a hypothesis generator, not a diagnostician -- every result required ACMG/AMP expert review, additional testing, and CLIA-certified lab confirmation before being returned to families.
- Key capabilities included inferring structural variants not in the input data, proposing digenic (two-gene) explanations, and generating novel mechanistic hypotheses like a new S1PR1-vitiligo connection.
- The OpenAI Foundation is funding the Manton Center to build a platform-agnostic, low-cost genetics AI copilot to scale this reanalysis approach more broadly.
- Limitations include no blinding to model confidence scores, no measurement of time or cost savings, and no evaluation of structural variants or repeat expansions.
For families with children suffering from rare genetic diseases, the wait for a diagnosis can stretch across years -- sometimes decades. A new peer-reviewed study published in NEJM AI shows that AI can meaningfully chip away at that backlog. Researchers from Boston Children's Hospital's Manton Center for Orphan Disease Research, Harvard University, and OpenAI used o3 Deep Research to reanalyze 376 previously unsolved pediatric cases -- and surfaced leads that led to 18 confirmed diagnoses.
The Diagnostic Odyssey Problem
Rare disease diagnosis is one of medicine's hardest problems. Roughly half of patients with rare diseases remain undiagnosed after extensive testing and specialist review. The challenge is not just biological -- it is operational. A patient's phenotype descriptions, test results, and family history can be split across databases that use different identifiers, formats, and vocabularies. Experts may also sequence a child's genome before a relevant gene or its variants have been linked to disease.
This creates a compounding maintenance problem. The patient's genome may stay the same, but the evidence around it keeps changing: researchers link new genes and variants to disease, labs reclassify old variants, and case databases and papers accumulate new observations. Each update means old inconclusive cases are potentially worth revisiting -- but doing so at scale is nearly impossible with human effort alone.
How the Workflow Actually Ran
The team did not simply feed raw genomic data into a chatbot. For each case, they assembled a de-identified packet containing standardized Human Phenotype Ontology (HPO) terms -- a controlled vocabulary for describing symptoms -- along with clinician notes, patient metadata, and a filtered variant table capturing each variant's rarity, predicted protein effect, ClinVar classification, and signal quality across available family members. HPO terms are essentially a shared language that lets computers and clinicians describe symptoms consistently across databases.
The model was asked to propose the most plausible molecular explanation and to show its work. Researchers then reviewed outputs using the ACMG/AMP framework -- the standard clinical classification system for genetic variants -- with at least two team members reviewing each candidate. A finding only counted as a diagnosis after a CLIA-certified lab confirmed it and the clinical team returned the result to the family. The AI's job was hypothesis generation, not diagnosis.
Before touching unsolved cases, the team stress-tested the workflow on cases with known answers. It recovered the correct gene and variant in duplicate runs for 48 of 51 cases across a variety of rare conditions. In a set of 57 neuromuscular cases, the workflow returned the correct diagnosis in duplicate runs for 45 cases. The model's self-reported confidence scores also tracked meaningfully with accuracy: the mean minimum score was 85.6 for consistently correct calls and 42.1 for incorrect or unknown calls.
The Numbers
The team applied the workflow to four groups of previously unsolved cases: children with neurodevelopmental conditions, people with rare neuromuscular disease, children and adolescents with early psychosis, and cases of sudden unexpected death in pediatrics. Here is how the results broke down:
| Cohort | Cases | Diagnoses | Yield |
|---|---|---|---|
| Neurodevelopmental | 100 | 10 | 10.0% |
| Neuromuscular disease | 61 | 4 | 6.6% |
| Sudden unexpected death (pediatric) | 200 | 2 | 1.0% |
| Early psychosis | 15 | 2 | 13.3% |
| Total | 376 | 18 | 4.8% |
A 4.8% yield sounds modest, but context matters enormously here. A meta-analysis of 29 studies comprising 9,419 undiagnosed patients reported an average increase in diagnostic yield of just 10% after a median of about 24 months of standard reanalysis. The o3 workflow achieved its gains on cases that had already been through that kind of expert review. One independent expert called a 5% diagnostic yield "truly meaningful" and said it "could serve as a significant screening tool to help speed up" diagnosis.
What the Model Actually Did Differently
Several findings illustrate why a general-purpose reasoning model adds something that traditional pipelines miss. In one early-psychosis case, the model inferred a structural event in the genome that was not listed in the input data. It connected a run of low-quality calls on chromosome 22 with the child's cardiac, immune, neurodevelopmental, and psychiatric features, then hypothesized a 22q11.2 deletion associated with DiGeorge syndrome -- which was confirmed with follow-up sequencing.
The model also went beyond the single-gene assumption baked into most diagnostic pipelines.
Although the prompt asked for one monogenic cause, the model sometimes surfaced two genes that better explained a complex presentation -- variants in LAMA2 and FOXP1 together helped account for muscle and neurodevelopmental features in one case; another had a previously unrecognized digenic explanation involving TTN and SRPK3.
In one neurodevelopmental case, the model went further still -- generating a hypothesis about a completely new biological mechanism.
It highlighted an 11-amino-acid deletion in S1PR1 in a person with vitiligo, integrating evidence suggesting the deletion could alter receptor structure and signaling in ways that reduce pigment production while helping immune cells persist in the skin.
That proposed relationship requires experimental validation, but it shows the model doing something closer to scientific reasoning than pattern matching.
Behind the Numbers: Kyra's Story
One of the four neuromuscular diagnoses belonged to a patient named Kyra.
Her case began at age 9, when her mother noticed she was not getting as low in her karate stances as she used to. What followed was a nearly 20-year journey through tests, treatments, and consultations without a diagnosis.
The team linked her condition to a frameshift variant in HSPB8 and diagnosed a form of myofibrillar myopathy -- a disease where abnormal protein structures build up in muscle fibers and contribute to weakness. A genetic counselor called Kyra about a week before her 28th birthday.
Who Wins -- and What the Limitations Are
The clearest winners are families stuck in diagnostic limbo. The Manton Center works with over 3,500 people affected by rare diseases, partnering with hospitals and health centers around the world. The study suggests that a growing backlog of unresolved genomic cases could be systematically revisited as medical knowledge advances, rather than waiting for a specialist to happen upon a new paper.
The study is careful about what it does not show. Researchers were not blinded to model confidence scores, the cohorts were heterogeneous, and the team did not measure time saved, cost, or false-positive workload. The study also did not systematically evaluate other forms of genetic variation such as structural variants, repeat expansions, deep-intronic changes, or mosaicism. Those are real gaps that prospective, multi-center trials will need to fill.
For the broader genomics field, the competitive picture shifts slightly. Purpose-built genomic AI tools have existed for years, but this study demonstrates that a general-purpose reasoning model -- without domain-specific fine-tuning -- can contribute meaningfully to one of the hardest diagnostic workflows in medicine. That raises the bar for specialized tools to demonstrate what they add beyond what a well-prompted frontier model can already do.
What Comes Next
The research does not stop here. The Manton Center is now building a platform-agnostic, low-cost genetics AI copilot that helps clinical teams reanalyze rare disease cases more quickly and consistently, with a focus on disease gene discovery, democratizing access to highly specialized diagnostics, and reducing inequities. This work is supported through the OpenAI Foundation's $50 million People-First AI Fund.
The study also points toward newer tools: purpose-built systems such as GPT-Rosalind are designed for deeper life-sciences work, including variant effects on protein structure and function -- capabilities not tested here that will require their own evaluations. The longer arc is toward making expert-level periodic reanalysis a routine part of rare disease care, not a one-time event. For thousands of families, the genome their child already had sequenced may already contain the answer -- it just needs the right reasoning layer to find it.
As Alan Beggs, director of the Manton Center, put it: "Researchers like Catherine and me can't possibly keep 8,000 different diseases in our heads. That's the power of AI."