Back

Pre-processing annotated homologous regions in protein sequences concerning machine-learning applications

Malhis, N.

2024-10-25 bioinformatics
10.1101/2024.10.25.620288 bioRxiv
Show abstract

Accurate preprocessing of annotated protein sequences with regard to homologies is essential for maintaining the integrity of machine-learning applications. This study presents two new tools--HAM (Homology-based Annotation Masking) and HAC (Homology Annotation Conflict)-- designed to address these challenges. HAM detects and masks homologous regions between datasets to prevent leakage, while HAC identifies and resolves annotation inconsistencies within datasets. Applying these tools to three benchmark datasets revealed substantial overlooked homology and annotation conflicts, even in datasets that had been previously clustered by sequence identity. These findings underscore the importance of homology-aware preprocessing to ensure the integrity of model training and evaluation. By integrating HAM and HAC into machine learning workflows, researchers can improve the consistency and trustworthiness of protein sequence-based predictions. Availabilitygithub.com/NawarMalhis/HAM.git

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.