Accuracy of somatic variant detection workflows for whole genome sequencing experiments
Jaksik, R.; Rosiak, J.; Zawadzki, P.; Sztromwasser, P.
Show abstract
Whole genome sequencing (WGS) becomes increasingly important for advancing personalized cancer care, driving not only basic science studies but also entering into clinical applications. Translating raw WGS data into the right clinical decision requires high accuracy of somatic variant detection, therefore novel data analysis methods have to be carefully evaluated. In this work we tested the performance of well-established somatic variant detection workflows: GATK, CPG-WGS, DRAGEN and Strelka2. By utilizing both real data, with well-defined mutations, and synthetic mutations spiked-in into real data, we were able to assess sensitivity and precision of each workflow, for various coverage and tumor purity levels. Individual tools excelled in different evaluation approaches, however the results demonstrated that DRAGEN has the highest overall performance when sensitivity is preferred over precision, and the opposite is true for CGP-WGS. The differences in results obtained using synthetic and real datasets, indicate that benchmarks based only on a single reference set may provide an incomplete picture.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Katdetectr: utilising unsupervised changepoint analysis for robust kataegis detection 96%
- cfDNA UniFlow: A unified preprocessing pipeline for cell-free DNA data from liquid biopsies 96%
- Genetic demultiplexing of pooled single-cell RNA-sequencing samples in cancer facilitates effective experimental design 95%
Similar papers in this journal
Similar papers in this journal
- iCOMIC: a graphical interface-driven bioinformatics pipeline for analyzing cancer omics data 95%
- Scalable and efficient DNA sequencing analysis on different compute infrastructures aiding variant discovery 95%
- Whole genome sequencing with AVITI and NovaSeq X Plus reveals comparable performance with contextual biases 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.