Quantifying the seed sensitivity of cancer subclonal reconstruction algorithms
Steinberg, P. L.; Liu, L. Y.; Neiman-Golden, A.; Patel, Y.; Boutros, P. C.
Show abstract
BackgroundIntra-tumoural heterogeneity complicates cancer prognosis and impairs treatment success. One of the ways subclonal reconstruction (SRC) quantifies intra-tumoural heterogeneity is by estimating the number of subclones present in bulk DNA sequencing data. SRC algorithms are probabilistic and need to be initialized by a random seed. However, the seeds used in bioinformatics algorithms are rarely reported in the literature. Thus, the impact of the initializing seed on SRC solutions has not been studied. To address this gap, we generated a set of ten random seeds to systematically benchmark the seed sensitivity of three probabilistic SRC algorithms: PyClone-VI, DPClust, and PhyloWGS. ResultsWe characterized the seed sensitivity of three algorithms across fourteen whole-genome sequences of head and neck squamous cell carcinoma and nine SRC pipelines, each composed of a single nucleotide variant caller, a copy number aberration caller and an SRC algorithm. This led to a total of 1470 subclonal reconstructions, including 1260 single-region and 210 multi-region reconstructions. The number of subclones estimated per patient vary across SRC pipelines, but all three SRC algorithms show substantial seed sensitivity: subclone estimates vary across different seeds for the same set of input using the same SRC algorithm. No seed consistently estimated the mode number of subclones across all patients for any SRC algorithm. ConclusionsThese findings highlight the variability in quantifying intra-tumoural heterogeneity introduced by the seed sensitivity of probabilistic SRC algorithms. We recommend that authors, reviewers and editors adopt guidelines to both report and randomize seed choices. It may also be valuable to consider seed-sensitivity in the benchmarking of newly developed SRC algorithms. These findings may be of interest in other areas of bioinformatics where seeded probabilistic algorithms are used and suggest consideration of formal seed reporting standards to enhance reproducibility.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- pyCancerSig: subclassifying human cancer with comprehensive single nucleotide, structural and microsatellite mutational signature deconstruction from whole genome sequencing 94%
- Generating realistic null hypothesis of cancer mutational landscapes using SigProfilerSimulator 94%
- Validation of genetic variants from NGS data using Deep Convolutional Neural Networks 94%
Similar papers in this journal
- Variant calling tool evaluation for variable size indel calling from next generation whole genome and targeted sequencing data 94%
- VarSCAT: A computational tool for sequence context annotations of genomic variants 93%
- eVIP2: Expression-based variant impact phenotyping to predict the function of gene variants 92%
Similar papers in this journal
Similar papers in this journal
- Assessing reliability of intra-tumor heterogeneity estimates from single sample whole exome sequencing data 95%
- Exogene: A performant workflow for detecting viral integrations from paired-end next-generation sequencing data 92%
- Curated Multiple Sequence Alignment for the Adenomatous Polyposis Coli (APC) Gene and Accuracy of In Silico Pathogenicity Predictions 92%
Similar papers in this journal
- Panel Informativity Optimizer (PIO): an R package to improve cancer NGS panel informativity 94%
- PanelCAT: an Open-Source Comparative Analysis Tool for Next-Generation Sequencing Panel Target Regions 93%
- Evaluating discordant somatic calls across mutation discovery approaches to minimize false negative drug-resistant findings 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.