Using somatic data to aid germline clinical variant interpretation in developmental disorders
Andrews, K. A.; Neville, M. D.; Martincorena, I.; Rahbari, R.; Firth, H.; Lindsay, S. J.; Tischkowitz, M.; Hurles, M.
Show abstract
Accurate interpretation of rare germline variants remains a major challenge in developmental disorders (DD). Somatic mutation data represent a largely untapped source of evidence for germline variant classi-fication. Identical or nearby mutations that drive positive selection when present in somatic tissues can cause developmental disorders when present in the germline. We integrated somatic mutation data from the Catalogue Of Somatic Mutations In Cancer (COSMIC), and healthy tissues (sperm and buccal epithelium) with germline variant datasets from ClinVar and large studies of de novo mutations in DD patients. Across 970 dominant DD genes, 195 have evidence of somatic selection, with a majority demonstrating concordant mechanisms between germline and somatic contexts. We benchmark the ability of somatic data to discriminate pathogenic from benign germline missense variation across dominant DD genes, identifying 145 genes in which somatic data are informative. The strongest utility is in altered-function genes where germline and somatic mechanisms are concordant, for example the RASopathy genes. In these genes, codon-level aggregation of somatic missense counts yields predictive performance comparable to computational predictors or MAVE assays (AUC-ROC 0.895 for somatic data, versus 0.893 for REVEL). Combining somatic features with computational scores improves discrimination further. Using likelihood ratios, we map COSMIC missense codon count thresholds onto American College of Medical Genetics and Genomics/Association for Molecular Pathology (ACMG/AMP)-style evidence strengths, showing that somatic data can reach strong levels of evidence in germline variant interpretation in DD and enable reclassification of variants of uncertain significance. Together, these results establish somatic mutation data as a scalable and clinically actionable evidence source for germline variant interpretation in select DD genes. Graphical abstract(Generated using FigureLabs) O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=104 SRC="FIGDIR/small/732808v1_ufig1.gif" ALT="Figure 1"> View larger version (40K): org.highwire.dtl.DTLVardef@150bec9org.highwire.dtl.DTLVardef@1dacf5org.highwire.dtl.DTLVardef@46121dorg.highwire.dtl.DTLVardef@4f5c38_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Estimating diagnostic noise in panel-based genomic analysis 96%
- Informing Variant Assessment using Structured Evidence from Prior Classifications (PS1, PM5, and PVS1 Sequence Variant Interpretation Criteria) 96%
- Assessment of the variant prioritisation strategy for genomic newborn screening in the Generation Study 96%
Similar papers in this journal
- Extracting and calibrating evidence of variant pathogenicity from population biobank data 96%
- Advanced variant classification framework reduces the false positive rate of predicted loss of function (pLoF) variants in population sequencing data 96%
- Availability of benign missense variant “truthsets” for validation of functional assays: current status and a novel systematic approach 95%
Similar papers in this journal
- Evaluating Genome Sequencing Strategies: Trio, Singleton, and Standard Testing in Rare Disease Diagnosis 95%
- Clustering of predicted loss-of-function variants in genes linked with monogenic disease can explain incomplete penetrance 94%
- A systematic analysis of splicing variants identifies new diagnoses in the 100,000 Genomes Project. 94%
Similar papers in this journal
- Assessing performance of pathogenicity predictors using clinically-relevant variant datasets 95%
- A comparative medical genomics approach may facilitate the interpretation of rare missense variation 93%
- The PS4-Likelihood Ratio Calculator: Flexible allocation of evidence weighting for case-control data in variant classification 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.