Back

Bridging functional annotation gaps in non-model plant genes with AlphaFold, DeepFRI and small molecule docking

Stephan, G.; Dugdale, B.; Deo, P.; Harding, R.; Dale, J.; Visendi, P.

2021-12-23 bioinformatics
10.1101/2021.12.22.473925 bioRxiv
Show abstract

BackgroundFunctional annotation assigns descriptive biological meaning to genetic sequences. Limited availability of manually curated or experimentally validated plant genes from a diverse range of taxa poses a significant challenge for functional annotation in non-model organisms. Accurate computational approaches are required. We argue that recent breakthroughs in deep learning have the potential to not only narrow the functional annotation gap between non-model and model plant organisms, but also annotate and reveal novel functions even for genes with no homologs in public databases. ResultsDeep learning models were applied to functionally annotate a set of previously published differentially expressed genes. Predicted protein structures and functional annotations were generated using the AlphaFold protein structure and DeepFRI protein language inference models respectively. The resulting structures and functional annotations were validated using small molecule docking experiments. DeepFRI and AlphaFold models not only correctly annotated differentially expressed genes, but also revealed detailed mechanisms involving protein-protein interactions. ConclusionsDeep learning models are capable of inferring novel functions and achieving high accuracy in functional annotation. Their increased use in plant research will result in major improvements in annotations for non-model plants that are underrepresented in genome databases. We illustrate how integrating protein structure prediction, functional residue prediction, and small molecule docking can infer plausible protein-protein interactions and yield additional mechanistic insights. This approach will aid in the selection of candidate genes for further study from differential expression studies that generate large gene lists.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Proteins: Structure, Function, and Bioinformatics
88 papers in training set
Top 0.1%
22.8%
2
Scientific Reports
3612 papers in training set
Top 0.5%
19.2%
3
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.4%
6.5%
4
ACS Omega
105 papers in training set
Top 0.4%
4.2%
50% of probability mass above
5
Protein Science
246 papers in training set
Top 1%
3.6%
6
Briefings in Bioinformatics
354 papers in training set
Top 2%
3.6%
7
Structure
193 papers in training set
Top 0.8%
2.9%
8
Journal of Biomolecular Structure and Dynamics
43 papers in training set
Top 0.5%
2.9%
9
Journal of Chemical Information and Modeling
238 papers in training set
Top 1%
2.5%
10
International Journal of Molecular Sciences
494 papers in training set
Top 6%
2.0%
11
PLOS Computational Biology
1863 papers in training set
Top 14%
1.8%
12
PLOS ONE
5266 papers in training set
Top 47%
1.8%
13
Communications Biology
993 papers in training set
Top 16%
1.6%
14
Bioinformatics
1204 papers in training set
Top 7%
1.2%
15
Frontiers in Molecular Biosciences
102 papers in training set
Top 1%
1.2%
16
The Journal of Physical Chemistry B
167 papers in training set
Top 1%
1.2%
17
Journal of Proteome Research
234 papers in training set
Top 1%
1.1%
18
BMC Bioinformatics
457 papers in training set
Top 5%
0.9%
19
Journal of Molecular Biology
232 papers in training set
Top 3%
0.9%
20
Frontiers in Bioinformatics
49 papers in training set
Top 2%
0.6%
21
Biomolecules
100 papers in training set
Top 3%
0.6%
22
Journal of Structural Biology
64 papers in training set
Top 0.8%
0.6%
23
F1000Research
88 papers in training set
Top 4%
0.6%
24
RSC Advances
22 papers in training set
Top 1%
0.5%
25
Bioinformatics Advances
203 papers in training set
Top 5%
0.5%
26
BMC Genomics
406 papers in training set
Top 10%
0.5%
27
International Journal of Biological Macromolecules
76 papers in training set
Top 3%
0.5%
28
iScience
1154 papers in training set
Top 43%
0.5%