Pre-processing annotated homologous regions in protein sequences concerning machine-learning applications
Malhis, N.
Show abstract
Accurate preprocessing of annotated protein sequences with regard to homologies is essential for maintaining the integrity of machine-learning applications. This study presents two new tools--HAM (Homology-based Annotation Masking) and HAC (Homology Annotation Conflict)-- designed to address these challenges. HAM detects and masks homologous regions between datasets to prevent leakage, while HAC identifies and resolves annotation inconsistencies within datasets. Applying these tools to three benchmark datasets revealed substantial overlooked homology and annotation conflicts, even in datasets that had been previously clustered by sequence identity. These findings underscore the importance of homology-aware preprocessing to ensure the integrity of model training and evaluation. By integrating HAM and HAC into machine learning workflows, researchers can improve the consistency and trustworthiness of protein sequence-based predictions. Availabilitygithub.com/NawarMalhis/HAM.git
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ProteinPrompt: a webserver for predicting protein-protein interactions 95%
- Reciprocal Best Structure Hits: Using AlphaFold models to discover distant homologues 95%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 94%
Similar papers in this journal
- DELPHI: accurate deep ensemble model for protein interaction sites prediction 96%
- NEFFy: A Versatile Tool for Computing the Number of Effective Sequences 96%
- Embedding-based alignment: combining protein language models and alignment approaches to detect structural similarities in the twilight-zone 96%
Similar papers in this journal
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 95%
- FlatProt: 2D visualization eases protein structure comparison 95%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 95%
Similar papers in this journal
- Tailored machine learning models for functional RNA detection in genome-wide screens 94%
- BRAKER2: Automatic Eukaryotic Genome Annotation with GeneMark-EP+ and AUGUSTUS Supported by a Protein Database 93%
- An Integrative Multitiered Computational Analysis for Better Understanding the Structure and Function of 85 Miniproteins 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.