Benchmarking datasets for machine learning in protein function prediction
Jang, Y.; Qin, Q.-Q.; Wang, J.-L.; Kornmann, B.
Show abstract
Remarkable progress has been achieved by machine learning, particularly in accurate prediction of protein tertiary structures. Despite these advances, accurately annotating protein functions through machine learning approaches remains challenging, primarily due to the limited availability of large-scale benchmarking data. In this study, we addressed this gap by systematically screening proteins from the UniProt database for functional annotations, resulting in the creation of a benchmarking dataset that includes protein sequences and their corresponding annotations. The Protein Annotation Dataset (PAD) is a resource available to train a wide range of machine learning models for assignment of function annotations to previously unlabeled proteins. We curated a comprehensive dataset comprising four categories of functional annotations using enzyme commission (EC) numbers and gene ontology (GO) terms. The dataset was subsequently partitioned into training, validation, and test subsets. Furthermore, we incorporated an independent set from 12 diverse species, enabling the development and evaluation of innovative machine learning models.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- PELSA-Decipher: a software tool for the processing and interpretation of ligand protein interaction dataset acquired by PELSA 93%
- Computational identification of human biological processes and protein sequence motifs putatively targeted by SARS-CoV-2 proteins using protein-protein interaction networks 92%
- DIMA: Data-driven selection of a suitable imputation algorithm 92%
Similar papers in this journal
Similar papers in this journal
- Designing of thermostable proteins with a desired melting temperature 95%
- PIPENN-EMB: ensemble net and protein embeddings generalise protein interface prediction beyond homology 95%
- Deep Learning for Protein Peptide bindingPrediction: Incorporating Sequence, Structural andLanguage Model Features 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.