Plant Bioengineering Atlas: A Knowledge Graph of Genes, DNA Constructs, and Plant Traits.
Yawar, K. A.; Martin, S.; Weston, D. J.; Gu, L.; Tuskan, G. A.; Yang, X.
Show abstract
Plant bioengineering has generated tens of thousands of genotype-to-phenotype relationships, but this knowledge remains fragmented across narrative literature and difficult to use computationally. Inconsistent descriptions of DNA constructs, host species, and traits, including variable species names, omitted regulatory elements, and inconsistent gene symbols, impede data reuse, comparative analysis, and design-build-test-learn cycles. Here, we present the Plant Bioengineering Atlas, a literature-mined, ontology-grounded knowledge base assembled using an artificial intelligence (AI)-aided extraction pipeline. A large language model parsed open-access primary research articles to generate structured, provenance-anchored records of engineered genes, modification types, promoter-gene-terminator constructs, host species, target traits, and reported phenotypes, with every record traceable to its source. The current release contains 14,358 curated records encompassing 6,998 distinct genes across 436 plant species from 6,452 papers published between 2000 and 2026. Corpus analysis reveals that experiments are concentrated in a small group of model and crop species, disease and pathogen resistance is the most frequently engineered trait class, and constitutive regulatory parts (particularly the CaMV 35S promoter and NOS terminator) remain pervasive. Two in five records omit one or both flanking regulatory elements (i.e., promoter and terminator), while only 23.4% describe cassettes in which both elements resolve to named part classes, exposing a systematic reproducibility gap. We organize these data into a knowledge graph linking genes, constructs, species, and traits; provide access through an interactive web portal; and propose an AI-compatible documentation standard for AI-ready reporting. The Plant Bioengineering Atlas provides a foundation for data-driven hypothesis generation and AI-aided plant biodesign.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- A near gap-free haplotype-resolved genome assembly of Zoysia japonica uncovers intra-subgenomic gene expression and regulatory variation 91%
- Jan and mini-Jan, a model system for potato functional genomics 90%
- CuBe: a geminivirus-based copper-regulated expression system suitable for post-harvest activation. 90%
Similar papers in this journal
- Diversification of gene expression across extremophytes and stress-sensitive species in the Brassicaceae 92%
- A history-dependent integrase recorder of plant gene expression with single cell resolution 91%
- Exceptional subgenome stability and functional divergence in allotetraploid teff, the primary cereal crop in Ethiopia 90%
Similar papers in this journal
- CactEcoDB: Trait, spatial, environmental, phylogenetic and diversification data for the cactus family 92%
- The Leipzig Catalogue of Vascular Plants (LCVP) - An improved taxonomic reference list for all known vascular plants 90%
- A haplotype-phased male genome sequence of the stinging nettle, Urtica dioica ssp. dioica 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.