Genetic constraint at single amino acid resolution improves missense variant prioritisation and gene discovery
Zhang, X.; Theotokis, P. I.; Li, N.; the SHaRe Investigators, ; Wright, C.; Samocha, K. E.; Whiffin, N.; Ware, J. S.
Show abstract
The clinical impact of most germline missense variants in humans remains unknown. Genetic constraint identifies genomic regions under negative selection, where variations likely have functional impacts, but the spatial resolution of existing constraint metrics is limited. Here we present the Homologous Missense Constraint (HMC) score, which measures genetic constraint at quasi single amino-acid resolution by aggregating signals across protein homologues. We identify one million possible missense variants under strong negative selection. HMC precisely distinguishes pathogenic variants from benign variants for both early-onset and adult-onset disorders. It outperforms existing constraint metrics and pathogenicity meta-predictors in prioritising de novo mutations from probands with developmental disorders (DD), and is orthogonal to these, adding power when used in combination. We demonstrate utility for gene discovery by identifying seven genes newly-significant associated with DD that could act through an altered-function mechanism. Overall, HMC is a novel and strong predictor to improve missense variant interpretation.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genome-wide prediction of dominant and recessive neurodevelopmental disorder risk genes 97%
- Detecting cryptic clinically-relevant structural variation in exome sequencing data increases diagnostic yield for developmental disorders 96%
- The landscape of autosomal-recessive pathogenic variants in European populations reveals phenotype-specific effects 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.