Searching sequence databases for functional homologs using profile HMMs: how to set bit score thresholds?
Srivastava, J.; Hembrom, R.; Kumawat, A.; Balaji, P. V.
Show abstract
MotivationUniProt and BFD databases together have 2.5 billion protein sequences. A large majority of these proteins have been electronically annotated. Automated annotation pipelines, vis-a-vis manual curation, have the advantage of scale and speed but are fraught with relatively higher error rates. This is because sequence homology does not necessarily translate to functional homology, molecular function specification is hierarchic and not all functional families have the same amount of experimental data that one can exploit for annotation. Consequently, customization of annotation workflow is inevitable to minimize annotation errors. ResultsWe discuss possible ways of customizing the search of sequence databases for functional homologs using profile HMMs. Choosing an optimal bit score threshold is a critical step in the application of HMMs, which is illustrated using four Case Studies; the single domain nucleotide sugar 6-dehydrogenase and lysozyme-C families, and SH3 and GT-A domains which are typically found as a part of multi-domain proteins. We also discuss the limitations of using profile HMMs for functional annotation and suggests some possible ways to partially overcome such limitations. Supplementary informationSupplementary_material containing Figures S1-S7 and Tables S1 and S2 Supplementary_dataset.xlsx
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Naegleria fowleri: protein structures to facilitate drug discovery for the deadly, pathogenic free-living amoeba 95%
- Phylogenetic analysis of the MCL1 BH3 binding groove and rBH3 sequence motifs in the p53 and INK4 protein families 94%
- Conserved intramolecular networks in GDAP1 are closely connected to CMT-linked mutations and protein stability 94%
Similar papers in this journal
- A Mathematical Genomics Perspective on the Moonlighting Role of Glyceraldehyde-3-Phosphate Dehydrogenase (GAPDH) 96%
- Proximal relationships of moonlighting Proteins in Escherichia coli: a mathematical genomic perspective 96%
- The Distal-Proximal Relationships Among the Human Moonlighting Proteins: Evolutionary hotspots and Darwinian checkpoints 96%
Similar papers in this journal
- Same, Same, but Different: Molecular Analyses of Streptococcus pneumoniae Immune Evasion Proteins Identifies new Domains and Reveals Structural Differences between PspC and Hic Variants 93%
- Predicting human and viral protein variants affecting COVID-19 susceptibility and repurposing therapeutics 93%
- In Silico Analysis Predicting Effects of Deleterious SNPs of Human RASSF5 Gene on its Structure and Functions 93%
Similar papers in this journal
Similar papers in this journal
- Application of a Machine Learning Approach Towards the Targeted Identification of Phage Depolymerases 94%
- HIHISIV: a database of gene expression in HIV and SIV host immune response 93%
- CysPresso: A classification model utilizing deep learning protein representations to predict recombinant expression of cysteine-dense peptides 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.