SC-BIG: A Hierarchical Bayesian Model for Bulk-Informed Single Nucleotide Variant Calling in Single Cells
Schuette, D.; Kono, T. J. Y.; Schwarz, R. F.
Show abstract
Single-cell DNA sequencing (scDNA-seq) has emerged as a primary method for studying the evolution of cancer genomes and intra-tumor heterogeneity. However, despite technological advances, scDNA-seq remains noisy and is affected by amplification biases and allelic dropouts. Accurately determining the presence or absence of candidate somatic nucleotide variants (SNVs) in individual cancer cells therefore remains challenging. One strategy to alleviate this issue is to perform bulk whole-genome sequencing simultaneously with single-cell sequencing. To date, only few computational methods have been developed for bulk-informed detection of somatic SNVs in single cells, and existing methods do not adequately account for somatic copy-number alterations or clonal admixtures. We here present SC-BIG, a hierarchical Bayesian model that leverages bulk sequencing data from a representative tumor sample to improve SNV detection. SC-BIG propagates uncertainty across multiple biological parameters, including copy number alterations, sample purity, and SNV clonality. In a first step, the cancer cell fraction (CCF) of a SNV is jointly estimated from bulk and single-cell data. The CCF in turn then acts as a prior in the second inference step to calculate per-cell posterior probabilities for the presence of the SNV. We demonstrate that across simulated scenarios of varying CCFs, SC-BIG outperforms both naive thresholding and ProSolo, the only bulk-informed single-cell mutation caller described so far. Importantly, SC-BIG produces well-calibrated posterior probabilities that provide interpretable uncertainty quantification, enabling direct integration into downstream analyses such as phylogenetic reconstruction.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.