Seqrutinator: Non-Functional Homologue Sequence Scrutiny for the Generation of large Datatsets for Protein Superfamily Analysis
Amalfitano, A.; Stocchi, N.; Atencio, H. M.; Villarreal, F.; ten Have, A.
Show abstract
BackgroundIn recent years protein bioinformatics has resulted in many good algorithms for multiple sequence alignment (MSA) and phylogeny. Little attention has been paid to sequence selection whereas notably recently published complete proteomes often have many sequences that are partial or derive from pseudogenes. Not only do these sequences add noise to the MSA, phylogeny and other downstream computational analyses, they also instigate many errors in the processing of the MSAs and downstream analyses, including the phylogeny. ObjectiveThis work aims to provide and test an objective, automated but flexible pipeline for the scrutiny of sequence sets from large, complex, eukaryotic protein superfamilies. The pipeline should classify sequences with high precision and recall as either functional or non-functional. The pipeline should classify no or only a few SwissProt sequences as non-functional (high precision) and sequences from other related superfamilies as non-functional (high recall) and result in a demonstrably much improved MSA (high performance). ResultsSeqrutinator is a pipeline that consists of five modules written in Python3 that identify and remove sequences that are likely Non-Functional Homologues (NFH). Here we tested the pipeline using three complex plant superfamilies (BAHD, CYP and UGT) that act in specialized metabolism, using the complete proteomes of 16 plant species as input and SwissProt as a control. Only 1.94% of SwissProt sequences with wetlab evidence were identified as NFH and all sequences from other related superfamilies were removed. Most NFH sequences are partial but, interestingly, their removal results in highly improved MSAs. a few but significant sequences that instigate large gaps were found. The five modules show similar behaviour when applied to the 16 sequence sets of the three analysed superfamilies. Pipelines with different module orders result in similar classifications and, moreover, show that different modules often detect the same sequences. Conclusion and perspectiveSeqrutinator forms a consistent pipeline for sequence scrutiny that does result in sequence sets that generate high fidelity MSAs. Recovery analyses show the method has high precision and recall.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- FINDER: An automated software package to annotate eukaryotic genes from RNA-Seq data and associated protein sequences 93%
- Plant Co-expression Annotation Resource: a webserver for identifying targets for genetically modified crop breeding pipelines 93%
- HH-suite3 for fast remote homology detection and deep protein annotation 93%
Similar papers in this journal
- DeNoFo: a file format and toolkit for standardised, comparable de novo gene annotation 95%
- GOThresher: a program to remove annotation biases from protein function annotation datasets 94%
- Embedding-based alignment: combining protein language models and alignment approaches to detect structural similarities in the twilight-zone 94%
Similar papers in this journal
- AvP: a software package for automatic phylogenetic detection of candidate horizontal gene transfers. 94%
- ECOD domain classification of 48 whole proteomes from AlphaFold Structure Database using DPAM 93%
- Towards a comprehensive view of the pocketome universe - biological implications and algorithmic challenges. 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.