Living Models' BOTANIC-1 Pinpoints One Melon Mutation out of 2,494 Candidates

A modular pipeline pairing Gemma 4 E4B with a plant genomic model ranked the true causal melon mutation #1 out of 2,494 candidates.

·
·
·
Read6 min
TypeNews
TopicLlms · Data
  • Living Models paired Gemma 4 E4B with BOTANIC-1, a plant genomic foundation model, as a modular agent pipeline.
  • Gemma orchestrates bioinformatics tool calls; BOTANIC-1 scores DNA variants using evolutionary constraint across 128 kb context.
  • In a melon test, the system ranked the true causal CmEIN3 mutation #1 out of 2,494 candidates.
  • Recall@1 jumped from 0.15 with classical tools to 0.90 with BOTANIC-1 across 20 runs, zero hallucinations.
  • On 500+ validated plant mutations, it hit top 0.1% in 15.9% of cases versus 4.6% for best baseline.
  • Runs locally on one NVIDIA L4 GPU via Ollama; BOTANIC-1 and demo are open.

A local AI stack ranks a melon mutation first

Paris-based biology lab Living Models has released a local workflow that combines Gemma 3n E4B with BOTANIC-1, a genomic language model trained on DNA from 320 plant species. In a preprint case study, the system ranked a validated melon mutation first among 2,494 single nucleotide candidates, with no tie.

Variant ranking sits between genetic mapping and laboratory validation, where researchers may face thousands of linked candidates that statistical tools cannot separate. A more accurate ranking can reduce the number of CRISPR constructs, greenhouse trials, and seasonal crosses required for confirmation. The authors describe the software stage as an afternoon of compute; biological validation still follows plant growth cycles.

When linked variants look identical

Plant breeders often use Bulk Segregant Analysis, or BSA, to associate a trait with a chromosomal region. Researchers group plants by phenotype, sequence their genomes, and search for variants that consistently travel with the trait.

During reproduction, neighboring stretches of DNA tend to be inherited together. A causal mutation can therefore share the same statistical pattern as hundreds or thousands of harmless variants in the surrounding block. When every candidate co-segregates across the sampled plants, correlation-based software lacks enough information to identify the causal base.

This condition, known as linkage disequilibrium, is usually resolved by growing more plants to capture rare recombination events or by testing individual candidates through gene editing. Both approaches can add months or years.

Two models, separate jobs

Living Models assigns orchestration to Gemma and sequence scoring to BOTANIC-1. Gemma writes shell pipelines, invokes bioinformatics tools, reads terminal output, and retries failed commands. BOTANIC-1 evaluates whether a nucleotide substitution is compatible with sequence patterns learned across plant genomes.

  • Gemma 3n E4B: Handles planning, coding, tool calls, file processing, and error recovery.
  • BOTANIC-1: Scores plant DNA variants using as much as roughly 128 kilobases of surrounding sequence.
  • Classical tools: Normalize, filter, and annotate variants with utilities such as bcftools and Ensembl VEP.

The modular design preserves Gemma’s general coding and tool-use abilities while giving raw nucleotide analysis to a purpose-built model. General text tokenizers represent DNA inefficiently, and long context windows alone do not encode regulatory or evolutionary constraints. Specialized genomic models supply that domain-specific signal.

The authors report running the stack locally through Ollama on a single NVIDIA L4 GPU. Local execution allows breeding programs to keep proprietary reference genomes and variant files within their own infrastructure.

The melon test narrows 47,492 variants

The case study targeted a known mutation in the melon CmEIN3 gene, which affects flower sex and fruit yield. Standard mapping reduced 47,492 genome-wide variants to 3,061 candidates on chromosome 2. The evaluation then compared three configurations:

  • Unguided Gemma: The model attempted to identify the causal variant without an expert-defined pipeline.
  • Expert-guided classical workflow: Gemma operated standard command-line bioinformatics tools.
  • Gemma with BOTANIC-1: Gemma prepared the candidates and sent eligible substitutions to the genomic model.

Across 16 tool-error events in the classical workflow, Gemma corrected nine after parsing the terminal output. One failure involved applying human-genome flags to plant annotation files. Even successful runs reached the statistical limit of the data: two coding mutations remained tied, including the validated mutation and an unrelated candidate.

Evolution supplies the tie-breaker

BOTANIC-1 calculates a log-likelihood score equivalent to log P(alternate | context) - log P(reference | context). A strongly negative value means the alternate nucleotide is much less compatible with the surrounding sequence patterns learned during training. That signal reflects evolutionary constraint, which provides evidence separate from the inheritance correlations measured by BSA.

The combined workflow filtered out 567 insertions, deletions, and rearrangements, leaving 2,494 single nucleotide variants. Gemma routed those substitutions to BOTANIC-1, and the validated CmEIN3 mutation ranked No. 1 without ties.

Repeated runs expose the gap

Results across 20 runs for each configuration
Configuration Recall@1 Runs with fabricated variants
Unguided Gemma 0% 3 of 20
Expert-guided classical workflow 15% 4 of 20
Gemma with BOTANIC-1 90% 0 of 20

Recall@1 measures how often the validated mutation appeared at the top of the ranking. “Fabricated variants” refers to reported candidates absent from the input variant set. The combined workflow placed the correct candidate first in 18 of 20 runs and produced no fabricated candidates.

Evidence beyond one melon

Across more than 500 published and experimentally validated plant mutations, BOTANIC-1 placed the causal mutation within the top 0.1% of candidates in 15.9% of cases. The strongest reported non-genomic-language-model baseline reached 4.6%. At the broader top-1% threshold, BOTANIC-1 reached 48.8%, compared with 33.9% for classical pipelines.

Retrospective ranking performance
Ranking threshold BOTANIC-1 Reported baseline
Top 0.1% 15.9% 4.6%
Top 1% 48.8% 33.9%

For a list of 2,500 variants, the top 0.1% contains roughly two or three candidates. These percentages measure how sharply the model can prioritize experiments; causal confirmation still requires genetic or molecular validation. The comparison suite included PlantCAD2 and Evo 2.

What is available

Workflow at a glance

  1. Use BSA or another mapping method to identify a candidate region.
  2. Export the reference sequence and candidate variants, typically in VCF format.
  3. Normalize, filter, and annotate the variants with standard bioinformatics tools.
  4. Send each eligible nucleotide substitution and its surrounding sequence to BOTANIC-1.
  5. Rank candidates by genomic model score and select the highest-priority variants for laboratory testing.

Keeping every model, tool, and input file on local hardware prevents sequence data from reaching a hosted API. Operators still need the usual controls around model weights, shell execution, file permissions, and reproducible environment versions.

Useful targets and firm limits

  • CRISPR prioritization: Rank constructs after BSA or quantitative trait locus mapping.
  • Historical datasets: Revisit candidate regions that remained unresolved because of linkage disequilibrium.
  • Private breeding data: Analyze proprietary genomes within local infrastructure.
  • Scientific agents: Adapt the orchestration pattern to other specialist foundation models.

Limits of the evidence

  • Preprint status: Peer review remains pending.
  • Retrospective design: The melon mutation and the broader benchmark variants were already known, so prospective discovery performance remains unmeasured.
  • Variant scope: The melon workflow excluded 567 insertions, deletions, and rearrangements before BOTANIC-1 scoring. Its reported success applies to single nucleotide variants.
  • Biological confirmation: A model score can prioritize candidates but cannot establish molecular mechanism or phenotype by itself.
  • Species coverage: Performance may vary for genomes and regulatory patterns poorly represented among the 320 training species.

A reusable scientific-agent pattern

The architecture gives each component a narrow responsibility: a compact language model manages files, commands, retries, and explanations, while a specialist model supplies the biological prior. The melon benchmark shows that this pairing can improve candidate ranking while keeping data local. Prospective studies will determine how much field, greenhouse, and gene-editing work it can save in active breeding programs.

Trending
  • No trending articles

Comments

avatar

Next Reads