ET-Pfam: Ensemble transfer learning for protein family prediction
Escudero, S.; Duarte, S.; Vitale, R.; Fenoy, E.; Bugnon, L. A.; Milone, D. H.; Stegmayer, G.
Show abstract
MotivationDue to the rapid growth of sequence generation, which has surpassed the expert curators ability to manually review and annotate them, the computational annotation of proteins remains a significant challenge in bioinformatics nowadays. The Pfam database contains a large collection of proteins that are nowadays annotated with domain families through multiple sequence alignments and profile Hidden Markov models (pHMMs). However, such computational annotation methods have some limitations such as problems for handling large datasets and the fact that multiple sequence alignments are computationally challenging to compute with high accuracy due to the increase in complexity as the number of sequences and lengths grow. Additionally, each HMM is independently obtained for each family missing the opportunity of learning patterns across families, that is from a complete view of all the dataset. As an alternative, some deep learning (DL) models have been recently proposed, nevertheless with simple representations of the inputs and moderate improvements in performance. ResultsIn this work we present ET-Pfam, a novel approach based on transfer learning and ensembles of multiple DL classifiers to predict functional families in the Pfam database. Several base DL models are first trained using learned representations from a protein large language model, with different hyperparameters to increase diversity. Then, the base models are integrated using classical ensemble strategies and novel voting approaches by learning weights for each model and for each Pfam family. Results demonstrate that the proposed ET-Pfam method can consistently diminish classification error rates compared to individual DL models, boosting prediction performance. Among the novel ensemble strategies presented here, the learned weights by family voting achieved the best performance, with the lowest error rate (7.00%), significantly surpassing the best individual base model error (12.91%) and three competitors of the state- of-the-art on the same Pfam dataset. AvailabilityData and source code are available at https://github.com/sinc-lab/ET-Pfam.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- DPCfam: a new method for unsupervised protein family classification 96%
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 95%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 95%
Similar papers in this journal
Similar papers in this journal
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 95%
- MULAN: Multimodal Protein Language Model for Sequence and Structure Encoding 95%
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.