An optimized variant prioritization process for rare disease diagnostics: recommendations for Exomiser and Genomiser
Cooperstein, I. B.; Marwaha, S.; Ward, A.; Kobren, S. N.; Carter, J. N.; Undiagnosed Diseases Network, ; Wheeler, M. T.; Marth, G. T.
Show abstract
PurposeWhole-exome sequencing (WES) and whole-genome sequencing (WGS) are increasingly used as standard genetic tests to identify the diagnostic variants in rare disease cases. However, prioritizing these variants to reduce the time and burden of manual interpretation by clinical teams remains a significant challenge. The Exomiser/Genomiser software suite is the most widely adopted open-source software for prioritizing coding and non-coding variants. Despite its ubiquitous use, limited data-driven guidelines currently exist to optimize its performance for diagnostic variant prioritization. Based on detailed analyses of Undiagnosed Diseases Network (UDN) probands, this study presents optimized parameters and practical recommendations for deploying the Exomiser and Genomiser tools. We also highlight scenarios where diagnostic variants may be missed and propose alternative workflows to improve diagnostic success in such complex cases. MethodsWe analyzed 386 diagnosed probands from the UDN, including cases with coding and non-coding diagnostic variants. We systematically evaluated how tool performance was affected by key parameters, including gene:phenotype association data, variant pathogenicity predictors, phenotype term quality and quantity, and the inclusion and accuracy of family variant data. ResultsParameter optimization significantly improved Exomisers performance over default parameters. For WGS data, the percentage of coding diagnostic variants ranked within the top ten candidates increased from 49.7% to 85.5%, and for WES, from 67.3% to 88.2%. For non-coding variants prioritized with Genomiser, the top ten rankings improved from 15.0% to 40.0%. We also explored refinement strategies for Exomiser outputs, including using p-value thresholds and flagging genes that are frequently ranked in the top 30 candidates but rarely associated with diagnoses. ConclusionThis study provides an evidence-based framework for variant prioritization in WES and WGS data using Exomiser and Genomiser. These recommendations have been implemented in the Mosaic platform to support the ongoing analysis of undiagnosed UDN participants and provide efficient, scalable reanalysis to improve diagnostic yield. Our work also highlights the importance of tracking solved cases and diagnostic variants that can be used to benchmark bioinformatics tools.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Advancing long-read nanopore genome assembly and accurate variant calling for rare disease detection 96%
- Extracting and calibrating evidence of variant pathogenicity from population biobank data 95%
- Advanced variant classification framework reduces the false positive rate of predicted loss of function (pLoF) variants in population sequencing data 95%
Similar papers in this journal
- The Importance of Automation in Genetic Diagnosis: Lessons from Analyzing an Inherited Retinal Degeneration Cohort with the Mendelian Analysis Toolkit (MATK) 96%
- Informing Variant Assessment using Structured Evidence from Prior Classifications (PS1, PM5, and PVS1 Sequence Variant Interpretation Criteria) 96%
- A gene pathogenicity tool 'GenePy' identifies missed biallelic diagnoses in the 100,000 Genomes Project 95%
Similar papers in this journal
- Evaluating Genome Sequencing Strategies: Trio, Singleton, and Standard Testing in Rare Disease Diagnosis 95%
- MetaRNN: Differentiating Rare Pathogenic and Rare Benign Missense SNVs and InDels Using Deep Learning 95%
- A systematic analysis of splicing variants identifies new diagnoses in the 100,000 Genomes Project. 94%
Similar papers in this journal
- Disease-specific prioritization of non-coding GWAS variants based on chromatin accessibility 95%
- Long-read genome sequencing for the diagnosis of neurodevelopmental disorders 95%
- IMPROVE-DD: Integrating Multiple Phenotype Resources Optimises Variant Evaluation in genetically determined Developmental Disorders 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.