Reconstructing intra-tumor fitness landscapes from scSeq CNA genotypes via simulation-based Bayesian inference and Deep Learning
KafiKang, M.; Skums, P.
Show abstract
Inferring the selective effects of copy-number alterations (CNAs) from clonal tumor data is essential for understanding tumor evolution. In practice, intra-tumor evolutionary parameters are typically estimated by fitting population genetic models to observed data using maximum likelihood or Bayesian methods. However, realistic mechanistic models often lead to intractable likelihoods, limiting the applicability of conventional inference approaches. Here, we introduce a likelihood-free, simulation-based framework for inferring intra-tumor selection coefficients directly from clonal CNA profiles. Our approach employs neural posterior estimation to amortize inference across simulated tumors and uses normalizing flows to flexibly parameterize high-dimensional posterior distributions while enabling robust uncertainty quantification. Our primary model, CloneMLP-NPE, learns representations of whole-tumor CNA genotypes using a multilayer perceptron (MLP)-based encoder. We compare this model against two baselines: (i) a Set Transformer encoder applied to the same whole-tumor CNA profiles, and (ii) a consensus-based approach that relies only on the CNA profile of the most abundant clone. On held-out simulations, CloneMLP-NPE achieves the strongest overall performance, yielding well-calibrated posterior distributions and more accurate posterior mean estimates than both baselines.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- PyClone-VI: Scalable inference of clonal population structures using whole genome data 95%
- SCONCE2: jointly inferring single cell copy number profiles and tumor evolutionary distances 94%
- eSVD-DE: Cohort-wide differential expression in single-cell RNA-seq data using exponential-family embeddings 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.