NCBoost v2: a classifier for non-coding variants in Mendelian diseases
Caron, B.; Rausell, A.
Show abstract
MotivationThe current diagnostic rate of rare diseases through whole-genome sequencing has stabilized at around 30% on average, highlighting the need for improved computational scores to identify pathogenic variants. In 2019, we developed NCBoost, a supervised-learning approach that mined a comprehensive set of sequence constraint features and proved particularly well suited to identifying high-effect pathogenic non-coding variants in genetic diseases. Since its first release, the substantial increase in the number of variants available for training, as well as the enhanced capacity to detect purifying selection signals from large-scale genome sequencing projects, motivated an update of NCBoost. ResultsWe implemented NCBoost v2, a pathogenicity score for non-coding single-nucleotide variants, trained on the largest set of curated pathogenic variants in monogenic Mendelian diseases available to date. It leverages conservation features computed from recent large-scale genomic consortia such as Zoonomia and gnomAD, and incorporates recent splice-altering predictive scores. NCBoost v2 outperformed alternative state-of-the-art methods in a variety of scenarios, providing more consistent scores across non-coding genomic regions and fine-tuning the scoring of pathogenic splice-altering variants in Mendelian disease genes. AvailabilityNCBoost v2 software is implemented in Python 3.10 and is freely available under the GNU General Public License Version 3 at https://doi.org/10.5281/zenodo.16029049 and https://github.com/RausellLab/NCBoost-2, together with precomputed scores for the human genome assembly GRCh38.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- varCADD: large sets of standing genetic variation enable genome-wide pathogenicity prediction 97%
- MetaRNN: Differentiating Rare Pathogenic and Rare Benign Missense SNVs and InDels Using Deep Learning 94%
- Genome-wide prediction of pathogenic gain- and loss-of-function variants from ensemble learning of diverse feature set 94%
Similar papers in this journal
Similar papers in this journal
- GeneBreaker: Variant simulation to improve the diagnosis of Mendelian rare genetic diseases 94%
- Matching whole genomes to rare genetic disorders: Identification of potential causative variants using phenotype-weighted knowledge in the CAGI SickKids5 clinical genomes challenge 94%
- Qatar Genome: Insights on Genomics from the Middle East 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.