Living Models pairs Gemma 4 with BOTANIC-1 to help decode plant DNA
By combining Gemma 4 with BOTANIC-1, a foundation model trained specifically on plant genomes, researchers could compress years of traditional crop genetics into hours of computation.
Modern crop breeding hinges on pinpointing causation: discovering why one row of melon vines collapses during drought while neighboring vines thrive, or how a single genetic switch alters floral architecture to multiply commercial yield. Tracking that flip across tens of thousands of candidate mutations has historically demanded years of field trials, seasonal plant crosses, and labor-intensive greenhouse screening.
To speed up this process, Paris-based AI biology lab Living Models tested a modular architecture pairing Gemma 4 with BOTANIC-1, a dedicated plant genomic world model. This pipeline demonstrates how modular architecture can untangle decades-old genetic puzzles in an afternoon and help agricultural scientists focus directly on high-confidence interventions.
Teaching Gemma to read DNA, without teaching it DNA
To isolate high-value traits, plant geneticists traditionally rely on Bulk Segregant Analysis—crossing opposing parent lines, cultivating the offspring across a growing season, sorting plants into trait-bearing pools, and sequencing their genomes. While this process narrows the trait to a broad chromosomal neighborhood, it leaves scientists with an unannotated spreadsheet containing tens of thousands of candidate mutations, including letter swaps (Single Nucleotide Polymorphisms, or SNPs), insertions, and deletions.
Faced with this bottleneck, a natural question arises: why not simply train a generalist model like Gemma directly on raw DNA?
Faced with this bottleneck, a natural question arises: why not simply train a generalist model like Gemma directly on raw DNA?
Developing reliable genomic capabilities within a generalist LLM introduces trade-offs: text-oriented tokenization may be inefficient for nucleotide sequences, supporting long inputs does not inherently capture long-range regulatory interactions, and genomic fine-tuning risks eroding the model’s core reasoning, coding, and tool-use faculties. Conversely, dedicated genomic language models (GLMs) like BOTANIC-1 excel at sequence representations but are difficult for non-technical researchers to deploy, query, and integrate into daily lab workflows.
Living Models resolves this by pairing the two: Gemma 4 serves as a conversational and agentic abstraction layer, translating natural-language biological inquiries into structured queries against BOTANIC-1’s inference endpoints.
The best language models are brilliant reasoners, but they are blind to the raw genome. Text taught Gemma 4 how to formulate hypotheses, write pipelines, and recover from errors like an expert bioinformatician. Pairing it with BOTANIC-1 gives Gemma the eyes to see what the plant is actually saying—allowing it not just to talk about biology, but to actually do biology.
Gemma 4 E4B acts as the autonomous bioinformatic orchestrator—writing terminal pipelines and executing tool calls—while BOTANIC-1, a foundation model pre-trained across 320 plant species, evaluates deep evolutionary constraints. Benchmarking showed larger models added operational overhead without improving accuracy, so Living Models runs E4B locally via Ollama on a single NVIDIA L4 GPU. This keeps proprietary genomes sovereign while the minimal footprint slashes operating costs, making domain-specialist AI accessible on standard commercial hardware.
Beyond correlation: The search for a causal signal
During plant reproduction, DNA is inherited in large, contiguous blocks. Harmless mutations hitchhike alongside the causal driver like passengers sharing the same train car. Because every variant in that block co-segregates identically across the sampled plants, statistical sorting software cannot distinguish which letter causes the trait and which are simply along for the ride.
Breaking these genetic ties can be time-consuming. Biologists might cultivate thousands of additional plants to capture rare recombination events, test other plant lines, or spend a year validating individual CRISPR constructs one by one.
To evaluate where AI creates an actual breakthrough in the lab, the Living Models team tested whether Gemma 4’s agentic reasoning could untangle a trait on its own using:
- Classical Tools (Setup A): Deploying Gemma as an autonomous bioinformatician using its agentic reasoning to write pipelines, filter coverage, and execute standard command-line software (bcftools, Ensembl VEP).
- BOTANIC-1 (Setup B): Equipping Gemma with BOTANIC-1, a domain-specific foundation model trained to read the evolutionary grammar of plant genomes.
The team turned to a landmark discovery led by collaborator Dr. Adnane Boualem, Research Director at INRAE, where a single nucleotide letter swap inside a melon’s CmEIN3 gene shifts flower sex from female to hermaphroditic—a decisive driver of pollination efficiency and fruit yield. Standard mapping narrowed the field from 47,492 variations genome-wide down to 3,061 regional candidates on Chromosome 2. Yet even on Chromosome 2, classical genetics hits an impassable physical wall: linkage disequilibrium.
Experiment A: Classical bioinformatics utilities hit a ceiling
Equipped with standard command-line tools, Gemma 4 operated with the procedural discipline of an autonomous computational biologist: formulating an execution plan, applying read-depth filters, and annotating mutations.
The model exhibited agentic resilience: when syntax errors occurred (such as attempting to apply human-genome flags to plant annotation files) Gemma parsed the terminal outputs and recovered autonomously in two out of three instances. Guided by filtration thresholds, in its best runs, Gemma narrowed the pool to two identical-scoring coding mutations. One was the validated CmEIN3 driver; the other was an unrelated gene.
To prove this deadlock was an inherent limitation of classical tools rather than the AI, the team executed the same data through deterministic scripts without the model:
While Gemma 4 successfully replicated the ceiling of human expert bioinformatics, that ceiling still leaves researchers at a standstill: because classical tools measure inheritance correlation rather than biological consequence, they cannot identify which mutation actually alters plant function—leaving breeders to spend months of lab validation on a coin toss.
Experiment B: BOTANIC-1 breaks the tie
In the second scenario, Gemma 4 was given a single specialized tool: the local BOTANIC-1 scoring function. Instead of relying on population frequency deltas, BOTANIC-1 evaluates variants across up to approximately 128 kb of sequence context using the Log-Likelihood Ratio (LLR) between the probabilities the model gives to the variant (alt) and to the original nucleotide in the reference genome (ref), given the surrounding genomic context (ctx) (Benegas, G. et al., PNAS, 2023):
This metric functions as an evolutionary constraint score; when a nucleotide position has remained conserved across millions of years of plant evolution, introducing an alternative base yields an intensely negative score indicating severe biological disruption.
Gemma 4 systematically parsed the Chromosome 2 candidates, filtered out 567 filtered out 567 insertions, deletions and rearrangements, and routed the remaining 2,494 single nucleotide mutations to the local BOTANIC-1 scoring function. The outcome broke the deadlock: the validated CmEIN3 mutation ranked #1 out of 2,494 candidates without ties.
Building this system means connecting two distinct capabilities: reasoning through an analysis and predicting the effects of changes in DNA. BOTANIC-1 learns from plant genomes to help distinguish variants that conventional analysis leaves tied. Gemma turns that signal into an actionable result: a prioritized hypothesis that a scientist can take to the lab.
By substituting population correlation with deep evolutionary constraint, BOTANIC-1 supplies the orthogonal biophysical signal needed to break through linkage disequilibrium. Equipping Gemma 4 with BOTANIC-1 raised Recall@1 from 0.00 (unguided) and 0.15 (expert‑guided) to 0.90 with no fabricated variants in 20 runs (against 3 and 4 of 20 with classical tools).
This capacity to shatter genetic ties extends far beyond a single melon chromosome. As detailed in Living Models’ research, the team evaluated BOTANIC-1 against classical bioinformatic tools and leading biological models—including PlantCAD2 and Evo2—across more than 500 published, experimentally validated causal plant mutations. Across this diverse retrospective benchmark, BOTANIC-1 placed the true causal mutation in the top 0.1% of candidates in 15.9% of cases, compared to just 4.5% for the best non-GLM baseline, and ranked the causal driver within the top 1.0% in 48.2% of cases (versus 32.7% for classical pipelines).
By grounding its analysis in deep evolutionary conservation, BOTANIC-1 can equip reasoning models with an empirical understanding of plant genomes that transcends the limits of raw correlation.
Explore the Gemma 4 x BOTANIC-1 traces to see how the experiment went in an interactive demo
A blueprint for scientific AI
This modular approach developed by Living Models establishes a repeatable blueprint for scientific AI. Rather than expecting a single model to master every layer of science, modern discovery benefits from a clean separation of concerns: generalist reasoning models like Gemma 4 excel at cognitive orchestration while specialized foundation models like BOTANIC-1 encode the physical and evolutionary constraints of nature. With lightweight models, agricultural labs and developers can execute these agentic pipelines directly on-premises, so intellectual property remains entirely sovereign and secure.
Genetics narrowed our search to a genomic region, but dozens of variants co-segregated identically with the trait. Separating them traditionally means growing and screening thousands of plants over months or years to capture rare recombination events. The real opportunity of AI is dramatically reducing that search space—identifying which constructs to validate first.
As accelerating climate shifts disrupt traditional growing seasons and challenge global food security, crop breeding can no longer afford to move at the pace of multi-year trial and error. By uniting open reasoning models with specialized genomic intelligence, researchers gain the speed, precision, and privacy needed to engineer resilient crops on a timeline the planet demands.

