Back

MmM: a machine learning approach to detect replication states and genomic subpop-ulations for single-cell DNA replication timing disentanglement

Josephides, J. M.; Chen, C.-L.

2023-12-26 genomics
10.1101/2023.12.26.573369 bioRxiv
Show abstract

We introduce MnM, an efficient tool for characterising single-cell DNA replication states and revealing genomic subpopulations in heterogeneous samples, notably cancers. MnM uses single-cell copy-number data to accurately perform missing-value imputation, classify cell replication states and detect genomic heterogeneity, which allows to separate somatic copy-number alterations from copy-number variations due to DNA replication. By applying our machine learning methods, our research unveils critical insights into chromosomal aberrations and showcases ubiquitous aneuploidy in tumorigenesis. MnM democratises single-cell subpopulation detection which, in hand, enables the extraction of single-cell DNA replication timing (scRT) profiles from genomically-heterogenous subpopulations detected by DNA content and issued from single samples. By analysing over 119,000 human single cells from cultured cell lines, patient tumours as well as patient-derived xenograft samples, the copy-number and replication timing profiles issued in this study lead to the first multi-sample subpopulation-disentangled scRT atlas and act as data contribution for further cancer research. Our results highlight the necessity of studying in vivo samples to comprehensively grasp the complexities of DNA replication, given that cell lines, while convenient, lack dynamic environmental factors. This tool offers to advance our understanding of cancer initiation and progression, facilitating further research in the interface of genomic instability and replication stress. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=150 SRC="FIGDIR/small/573369v2_ufig1.gif" ALT="Figure 1"> View larger version (30K): org.highwire.dtl.DTLVardef@25047forg.highwire.dtl.DTLVardef@4a57faorg.highwire.dtl.DTLVardef@d61dd5org.highwire.dtl.DTLVardef@140c6e0_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.