MissenseHMM: state-based annotations for missense variants through joint modeling of pathogenicity scores
Li, R.; Ernst, J.
Show abstract
Many computational predictors of missense variant pathogenicity are available. To capture information across various predictors, we propose MissenseHMM, which learns states corresponding to combinatorial patterns of variant prioritizations. We applied MissenseHMM to 43 predictors, annotating over 70 million missense variants with 20 states that showed distinct predictor scores patterns, amino acid substitutions and other genomic annotation enrichments. MissenseHMM state annotations enhanced individual predictors associations with clinical pathogenic variants and deep mutational scanning data, and also provided insight into the performances of various protein language models. Overall, MissenseHMM complements pathogenicity predictors and is an annotation resource for missense variant interpretation.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DeMAG predicts the effects of variants in clinically actionable genes by integrating structural and evolutionary epistatic features 97%
- PreMode predicts mode-of-action of missense variants by deep graph representation learning of protein sequence and structural context 96%
- G4mer: An RNA language model for transcriptome-wide identification of G-quadruplexes and disease variants from population-scale genetic data 96%
Similar papers in this journal
Similar papers in this journal
- Genomics 2 Proteins portal: A resource and discovery tool for linking genetic screening outputs to protein sequences and structures 96%
- Sliding Window INteraction Grammar (SWING): a generalized interaction language model for peptide and protein interactions 95%
- SCENIC+: single-cell multiomic inference of enhancers and gene regulatory networks 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.