A variant prioritization tool leveraging multiple instance learning for rare Mendelian disease genomic testing
Kim, H. H.; Baek, J. Y.; Han, H.; Jeong, W. C.; Kim, D.-W.; Kwon, K.; Song, Y.; Lee, H.; Seo, G. H.; Lee, J.; Lee, K.
Show abstract
BackgroundGenomic testing such as exome sequencing and genome sequencing is being widely utilized for diagnosing rare Mendelian disorders. Because of a large number of variants identified by these tests, interpreting the final list of variants and identifying the disease-causing variant even after filtering out likely benign variants could be labor-intensive and time-consuming. It becomes even more burdensome when various variant types such as structural variants need to be considered simultaneously with small variants. One way to accelerate the interpretation process is to have all variants accurately prioritized so that the most likely diagnostic variant(s) are clearly distinguished from the rest. MethodsTo comprehensively predict the genomic test results, we developed a deep learning based variant prioritization system that leverages multiple instance learning and feeds multiple variant types for variant prioritization. We additionally adopted learning to rank (LTR) for optimal prioritization. We retrospectively developed and validated the model with 5-fold cross-validation in 23,115 patients with suspected rare diseases who underwent whole exome sequencing. Furthermore, we conducted the ablation test to confirm the effectiveness of LTR and the importance of permutational features for model interpretation. We also compared the prioritization performance to publicly available variant prioritization tools. ResultsThe model showed an average AUROC of 0.92 for the genomic test results. Further, the model had a hit rate of 96.8% at 5 when prioritizing single nucleotide variants (SNVs)/small insertions and deletions (INDELs) and copy number variants (CNVs) together, and a hit rate of 95.0% at 5 when prioritizing CNVs alone. Our model outperformed publicly available variant prioritization tools for SNV/INDEL only. In addition, the ablation test showed that the model using LTR significantly outperformed the baseline model that does not use LTR in variant prioritization (p=0.007). ConclusionA deep learning model leveraging multiple instance learning precisely predicted genetic testing conclusion while prioritizing multiple types of variants. This model is expected to accelerate the variant interpretation process in finding the disease-causing variants more quickly for rare genetic diseases.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genome Alert!: a standardized procedure for genomic variant reinterpretation and automated genotype-phenotype reassessment in clinical routine 96%
- The Importance of Automation in Genetic Diagnosis: Lessons from Analyzing an Inherited Retinal Degeneration Cohort with the Mendelian Analysis Toolkit (MATK) 96%
- Reducing Sanger Confirmation Testing through False Positive Prediction Algorithms 95%
Similar papers in this journal
- Discordance between a deep learning model and clinical-grade variant pathogenicity classification in a rare disease cohort 95%
- Comprehensive reanalysis for CNVs in ES data from unsolved rare disease cases results in new diagnoses 95%
- WEGS: a cost-effective sequencing method for genetic studies combining high-depth whole exome and low-depth whole genome 92%
Similar papers in this journal
- MetaRNN: Differentiating Rare Pathogenic and Rare Benign Missense SNVs and InDels Using Deep Learning 97%
- Evaluating Genome Sequencing Strategies: Trio, Singleton, and Standard Testing in Rare Disease Diagnosis 95%
- Genome-Wide Sequencing as a First-Tier Screening Test for Short Tandem Repeat Expansions 94%
Similar papers in this journal
- Long-read genome sequencing for the diagnosis of neurodevelopmental disorders 94%
- IMPROVE-DD: Integrating Multiple Phenotype Resources Optimises Variant Evaluation in genetically determined Developmental Disorders 94%
- Omics-informed CNV calls reduce false positive rate and improve power for CNV-trait associations 94%
Similar papers in this journal
- Prioritization of disease genes from GWAS using ensemble based positive-unlabeled learning 94%
- Re-evaluation and Re-analysis of 152 research exomes five years after the initial report reveals clinically relevant changes in 20% 93%
- Exploiting Family History in Aggregation Unit-based Genetic Association Tests 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.