Back

The Case for Retaining Natural Language Descriptions of Phenotypes in Plant Databases and a Web Application as Proof of Concept

Braun, I. R.; Bassham, D. C.; Lawrence-Dill, C. J.

2021-02-06 bioinformatics
10.1101/2021.02.04.429796 bioRxiv
Show abstract

Similarities in phenotypic descriptions can be indicative of shared genetics, metabolism, and stress responses, to name a few. Finding and measuring similarity across descriptions of phenotype is not straightforward, with previous successes in computation requiring a great deal of expert data curation. Natural language processing of free text descriptions of phenotype is often less resource intensive than applying expert curation. It is therefore critical to understand the performance of natural language processing techniques for organizing and analyzing biological datasets and for enabling biological discovery. For predicting similar phenotypes, a wide variety of approaches from the natural language processing domain perform as well as curation-based methods. These computational approaches also show promise both for helping curators organize and work with large datasets and for enabling researchers to explore relationships among available phenotype descriptions. Here we generate networks of phenotype similarity and share a web application for querying a dataset of associated plant genes using these text mining approaches. Example situations and species for which application of these techniques is most useful are discussed. Database URLsThe database and analytical tool called QuOATS are available at https://quoats.dill-picl.org/. Code for the web application is available at https://git.io/Jtv9J. Datasets are available for direct access via https://zenodo.org/record/7947342#.ZGwAKOzMK3I. The code for the analyses performed for the publication is available at https://github.com/Dill-PICL/Plant-data and https://github.com/Dill-PICL/NLP-Plant-Phenotypes.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Applications in Plant Sciences
23 papers in training set
Top 0.1%
21.8%
2
Plant Physiology
238 papers in training set
Top 0.4%
12.6%
3
Bioinformatics
1204 papers in training set
Top 3%
7.2%
4
Plant Direct
95 papers in training set
Top 0.5%
5.4%
5
in silico Plants
27 papers in training set
Top 0.1%
4.8%
50% of probability mass above
6
G3 Genes|Genomes|Genetics
351 papers in training set
Top 1%
3.4%
7
PLOS Computational Biology
1863 papers in training set
Top 9%
3.4%
8
GENETICS
483 papers in training set
Top 2%
3.2%
9
Genome Biology
637 papers in training set
Top 5%
2.4%
10
Frontiers in Plant Science
256 papers in training set
Top 3%
2.4%
11
BMC Bioinformatics
457 papers in training set
Top 3%
2.1%
12
PeerJ
308 papers in training set
Top 5%
1.9%
13
The Plant Genome
57 papers in training set
Top 0.6%
1.7%
14
Bioinformatics Advances
203 papers in training set
Top 3%
1.4%
15
Frontiers in Genetics
230 papers in training set
Top 4%
1.1%
16
G3: Genes, Genomes, Genetics
252 papers in training set
Top 3%
1.1%
17
GigaScience
212 papers in training set
Top 3%
1.1%
18
Scientific Reports
3612 papers in training set
Top 66%
1.1%
19
Journal of Computational Biology
48 papers in training set
Top 0.8%
1.1%
20
PLOS ONE
5266 papers in training set
Top 55%
1.1%
21
BMC Genomics
406 papers in training set
Top 7%
1.0%
22
Frontiers in Bioinformatics
49 papers in training set
Top 1%
1.0%
23
Molecular Plant-Microbe Interactions®
57 papers in training set
Top 0.9%
1.0%
24
NAR Genomics and Bioinformatics
242 papers in training set
Top 4%
1.0%
25
New Phytologist
346 papers in training set
Top 5%
0.8%
26
Briefings in Bioinformatics
354 papers in training set
Top 7%
0.8%
27
The Plant Cell
161 papers in training set
Top 2%
0.6%