Perturbation robustness analyses reveal important parameters in variant interpretation pipelines
Wang, Y.; Adhikari, A. N.; Sunderam, U.; Kvale, M. N.; Currier, R. J.; Gallagher, R. C.; Kwok, P.-Y.; Puck, J. M.; Srinivasan, R.; Brenner, S. E.
Show abstract
MotivationGenome sequencing is being used routinely in clinical and research applications, but subsequent variant interpretation pipelines can vary widely. A systematic approach for exploring parameter choices and selection plays an important role in designing robust pipelines for specific clinical applications. ResultsWe present a framework to be applied in scenarios with limited data whereby expert knowledge informs pipeline refinement. Starting from initial reference variant interpretation pipelines with commonly used parameters, we derived pipelines by perturbing the parameters one by one to determine which parameters can yield meaningful changes in a pipelines performance. We updated the reference pipeline by fixing the value of parameters which have small impact on the pipelines performance. Then we conducted new rounds of perturbation as the process converged, yielding a stable pipeline which is robust. We applied the framework for genetic disease prediction in de-identified exomes from a cohort of 138 individuals with rare Mendelian inborn errors of metabolism (IEMs) and systematically explored how perturbing different parameters affected the pipelines sensitivity and specificity. For this application, we perturbed commonly used parameters in variant interpretation pipelines, including choices of genes, variant callers, transcript models, databases of allele frequencies, databases of curated disease variants, and tools for variant impact prediction. Our analyses showed that choice of variant callers, variant impact prediction tools, MAF threshold, and MAF databases can meaningfully alter results from a pipeline. This work informs the development of exome analysis pipelines designed for newborn metabolic disorder screening and suggests the general application of perturbation analysis in genome interpretation pipeline design.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Identifying digenic disease genes using machine learning in the undiagnosed diseases network 95%
- HiFi long-read genomes for difficult-to-detect clinically relevant variants 94%
- MRSD: a novel quantitative approach for assessing suitability of RNA-seq in the clinical investigation of mis-splicing in Mendelian disease 94%
Similar papers in this journal
- GeneBreaker: Variant simulation to improve the diagnosis of Mendelian rare genetic diseases 95%
- Phasing of de novo mutations using a scaled-up multiple amplicon long-read sequencing approach 95%
- Matching whole genomes to rare genetic disorders: Identification of potential causative variants using phenotype-weighted knowledge in the CAGI SickKids5 clinical genomes challenge 95%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.