Back

ATAClone: Cancer Clone Identification and Copy Number Estimation from Single-cell ATAC-seq

Cain, L. D.; Trigos, A. S.

2026-03-13 bioinformatics
10.64898/2026.03.11.710984 bioRxiv
Show abstract

Single-cell analyses of cancer typically begin by identifying distinct populations of cancer cells by unsupervised clustering. However, in many cases this clustering is explained simply by differences in DNA copy number, which affects the interpretation of differential expression results and tumour heterogeneity studies. To detect and estimate these differences in copy number, we have developed ATAClone. Applicable to both standalone and multiome scATAC-seq assays, ATAClone first identifies cancer cells with shared DNA copy number profiles (i.e. clones), then estimates their copy number jointly. Importantly, ATAClone can determine an optimal clustering resolution automatically using simulations. By utilising only stably accessible regions, ATAClone maximises copy number signal while minimising unrelated biological and technical noise. Additionally, by leveraging differences in total DNA between cells, ATAClone can infer absolute copy number, even in the presence of polyploidy. Using cancer cell mixture experiments, we verify the ability of ATAClone to accurately separate clones based on copy number differences. Moreover, using matched scATAC-seq and bulk whole genome sequencing, we show that copy number estimates from ATAClone are more accurate than those derived with existing methods, achieving Pearson correlations between 0.75-0.95 with their bulk-derived estimates. ATAClone represents an important tool for disentangling the genetic and non-genetic contributions to gene expression in cancer, providing deeper insight into the evolutionary history and adaptive forces driving a tumour.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.