CNVscore calculates pathogenicity scores for copy number variants together with uncertainty estimates accounting for learning biases in reference Mendelian disorder datasets.
Requena, F.; Salgado, D.; Malan, V.; Sanlaville, D.; Bilan, F.; Beroud, C.; Rausell, A.
Show abstract
Copy number variants (CNVs) are a major cause of rare pediatric diseases with a broad spectrum of phenotypes. Genetic diagnosis based on comparative genomic hybridization tests typically identifies [~]8-10% of patients as having CNVs of unknown significance, revealing the current limits of clinical interpretation. The adoption of whole-genome sequencing (WGS) as a first-line genetic test has significantly increased the load of CNVs identified in single genomes. Alongside short- and long-read sequencing technologies, a number of pathogenicity scores have been developed for filtering and prioritizing large sets of candidate CNVs in clinical settings. However, current approaches are often based, either explicitly or implicitly, on clinically annotated reference sets, which are likely to bias their predictions. In this study we developed CNVscore, a supervised-learning approach combining tree ensembles and a Bayesian classifier trained on pathogenic and non-pathogenic CNVs from reference databases. Unlike previous approaches, CNVscore couples pathogenicity estimates with uncertainty scores, making it possible to evaluate the suitability of a model for the query CNVs. Comprehensive comparative benchmark tests across independent sets and against alternative methods showed that CNVscore effectively distinguishes between pathogenic and benign CNVs. We also found that CNVs associated with CNVscores of low uncertainty were predicted with significantly higher accuracy than those of high uncertainty. However, the performance of current scoring approaches, including CNVscore, was compromised on CNV sets enriched in highly uncertain variants and presenting unconventional features, such as functionally relevant non-coding elements or the presence of disease genes irrelevant for the clinical phenotypes investigated. Finally, we used the CNVscore framework to guide CNV scoring model selection for the French National Database of Constitutional CNVs (BANCCO), which includes clinical diagnosis annotations. The CNVscore framework provides an objective strategy for leveraging the uncertainty on bioinformatic predictions to enhance the assessment of CNV pathogenicity in rare-disease cohorts. CNVscore is available as open-source software from https://github.com/RausellLab/CNVscore and is integrated into the CNVxplorer webserver http://cnvxplorer.com.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- MetaRNN: Differentiating Rare Pathogenic and Rare Benign Missense SNVs and InDels Using Deep Learning 94%
- Evaluating Genome Sequencing Strategies: Trio, Singleton, and Standard Testing in Rare Disease Diagnosis 93%
- Genome-Wide Sequencing as a First-Tier Screening Test for Short Tandem Repeat Expansions 93%
Similar papers in this journal
- Comprehensive reanalysis for CNVs in ES data from unsolved rare disease cases results in new diagnoses 95%
- Distinguishing benign from pathogenic duplications involving GPR101 and VGLL1-adjacent enhancers in the clinical setting with the bioinformatic tool POSTRE 93%
- Discordance between a deep learning model and clinical-grade variant pathogenicity classification in a rare disease cohort 92%
Similar papers in this journal
- The impact of 22q11.2 copy number variants on human traits in the general population 96%
- HiFi long-read genomes for difficult-to-detect clinically relevant variants 95%
- Exome copy number variant detection, analysis and classification in a large cohort of families with undiagnosed rare genetic disease 95%
Similar papers in this journal
- TADA - a Machine Learning Tool for Functional Annotation based Prioritisation of Putative Pathogenic CNVs 95%
- MINTIE: identifying novel structural and splice variants in transcriptomes using RNA-seq data 94%
- STRling: a k-mer counting approach that detects short tandem repeat expansions at known and novel loci 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.