Revisiting co-expression-based automated function prediction in yeast with neural networks and updated Gene Ontology annotations
McGuire, C. E.; Hibbs, M. A.
Show abstract
Automated function prediction (AFP) is the process of predicting the function of genes or proteins with machine learning models trained on high-throughput biological data. Deep learning with neural networks has become the dominant machine learning architecture of contemporary AFP models. However, it is unclear what difference exists between neural networks and previous machine learning architectures for AFP. Therefore, we created a model of AFP in yeast using neural networks that is trained on gene co-expression data to predict Gene Ontology (GO) labels. When trained on the same input data, we found that our model outperforms two other experimentally-validated co-expression-based AFP models using other machine learning techniques (Bayesian networks and adaptive query-driven search) when predicting individual genes involved in mitochondrion organization. In particular, we found our neural network model better distinguished mis-annotated negatives in its training data. Finally, we quantified how differences in the gene expression data and Gene Ontology annotations affect the performance of our model across each of its predicted GO terms. Our results suggest that neural networks are more performant and robust to GO mis-annotations compared to other machine learning architectures for co-expression-based AFP of some biological processes.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CoVar: A generalizable machine learning approach to identify the coordinated regulators driving variational gene expression 96%
- Reconstruction Set Test (RESET): a computationally efficient method for single sample gene set testing based on randomized reduced rank reconstruction error 95%
- Joint representation of molecular networks from multiple species improves gene classification 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.