Benchmarking Recent Computational Tools for DNA-binding Protein Identification
Luo, X.; Lin, A.; Chi, S.; Wong, L.; Rahman, C. R.
Show abstract
Identification of DNA-binding proteins (DBPs) is a crucial task in genome annotation, as it aids in understanding gene regulation, DNA replication, transcriptional control and various cellular processes. In this paper, we conduct an unbiased benchmarking of eleven state-of-the-art computational tools as well as traditional tools such as ScanProsite, BLAST, and HMMER for identifying DBPs. We highlight the data leakage issue in conventional datasets leading to inflated performance. We introduce new evaluation datasets to support further development. Through a comprehensive evaluation pipeline, we identify potential limitations in models, feature extraction techniques and training methods; and recommend solutions regarding these issues. We show that combining the predictions of the two best computational tools with BLAST based prediction significantly enhances DBP identification capability. We provide this consensus method as user-friendly software. The datasets and software are available at: https://github.com/Rafeed-bot/DNA_BP_Benchmarking. 1. Key PointsO_LIWe designed a comprehensive evaluation pipeline which systematically evaluates eleven recent machine learning (ML) based DBP identification tools. C_LIO_LIWe analyzed the test prediction mistakes made by top-performing tools identifying their potential limitations in terms of model architecture, feature extraction and class balancing. C_LIO_LIWe showed that although the best of these tools do not convincingly outperform BLAST, they still provide substantial value when integrated together with BLAST into a simple majority-voting ensemble. C_LIO_LIWe provide recommendations on more robust development & evaluation and better usability of future tools. C_LIO_LIWe provide the two best-performing ML-based tools, BLAST and the ensemble method as user-friendly software, as well as our proposed datasets, publicly available via GitHub. C_LI
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DELPHI: accurate deep ensemble model for protein interaction sites prediction 98%
- Embedding-based alignment: combining protein language models and alignment approaches to detect structural similarities in the twilight-zone 96%
- NEFFy: A Versatile Tool for Computing the Number of Effective Sequences 96%
Similar papers in this journal
Similar papers in this journal
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 96%
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 95%
- MULAN: Multimodal Protein Language Model for Sequence and Structure Encoding 95%
Similar papers in this journal
Similar papers in this journal
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 96%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 96%
- Binding affinity prediction for protein-ligand complex using deep attention mechanism based on intermolecular interactions 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.