MaCroDNA: Accurate integration of single-cell DNA and RNA data for a deeper understanding of tumor heterogeneity
Edrisi, M.; Huang, X.; Ogilvie, H.; Nakhleh, L.
Show abstract
Cancers develop and progress as mutations accumulate, and with the advent of single-cell DNA and RNA sequencing, researchers can observe these mutations, their transcriptomic effects, and predict proteomic changes with remarkable temporal and spatial precision. However, to connect genomic mutations with their transcriptomic and proteomic consequences, cells with either only DNA data or only RNA data must be mapped to a common domain. For this purpose, we present MaCroDNA, a novel method which uses maximum weighted bipartite matching of per-gene read counts from single-cell DNA and RNA-seq data. Using ground truth information from colorectal cancer data, we demonstrate the overwhelming advantage of MaCroDNA over existing methods in accuracy and speed. Exemplifying the utility of single-cell data integration in cancer research, we propose, based on results derived using MaCroDNA, that genomic mutations of large effect size increasingly contribute to differential expression between cells as Barretts esophagus progresses to esophageal cancer.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- scDesign2: a transparent simulator that generates high-fidelity single-cell gene expression count data with gene correlations captured 97%
- DelSIEVE: cell phylogeny model of single nucleotide variants and deletions from single-cell DNA sequencing data 97%
- Heterogeneous pseudobulk simulation enables realistic benchmarking of cell-type deconvolution methods 96%
Similar papers in this journal
Similar papers in this journal
- Hierarchical progressive learning of cell identities in single-cell data 97%
- Flexible Experimental Designs for Valid Single-cell RNA-sequencing Experiments Allowing Batch Effects Correction 96%
- Statistical modeling, estimation, and remediation of sample index hopping in multiplexed droplet-based single-cell RNA-seq data 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.