CellCover Defines Conserved Cell Types and Temporal Progression in scRNA-seq Data across Mammalian Neocortical Development
Ji, L.; Wang, A.; Sonthalia, S.; Naiman, D. Q.; Younes, L.; Colantuoni, C.; Geman, D.
Show abstract
1Definition of cell classes across the tissues of living organisms is central in the analysis of growing atlases of single-cell RNA sequencing (scRNA-seq) data across biomedicine. Marker genes for cell classes are most often defined by differential expression (DE) methods that serially assess individual genes across landscapes of diverse cells. This serial approach has been extremely useful, but is limited because it ignores possible redundancy or complementarity across genes that can only be captured by analyzing multiple genes simultaneously. Interrogating binarized expression data, we aim to identify discriminating panels of genes that are specific to, not only enriched in, individual cell types. To efficiently explore the vast space of possible marker panels, leverage the large number of cells often sequenced, and overcome zero-inflation in scRNA-seq data, we propose viewing marker gene panel selection as a variation of the "minimal set-covering problem" in combinatorial optimization. Using scRNA-seq data from blood and brain tissue, we show that this new method, CellCover, performs as good or better than DE and other methods in defining cell-type discriminating gene panels, while reducing gene redundancy and capturing cell-class-specific signals that are distinct from those defined by DE methods. Transfer learning experiments across mouse, primate, and human data demonstrate that CellCover identifies markers of conserved cell classes in neocortical neurogenesis, as well as developmental progression in both progenitors and neurons. Exploring markers of human outer radial glia (oRG, or basal RG) across mammals, we show that transcriptomic elements of this key cell type in the expansion of the human cortex likely appeared in gliogenic precursors of the rodent before the full program emerged in neurogenic cells of the primate lineage. We have assembled the public datasets we use in this report within the NeMO Analytics multi-omic data exploration environment [1], where the expression of individual genes (NeMO: Individual genes in cortex and NeMO: Individual genes in blood) and marker gene panels (NeMO: Telley 3 CellCover Panels, NeMO: Telley 12 CellCover Panels, NeMO: Sorted Brain Cell CellCover Panels, and NeMO: Blood 34 CellCover Panels) can be freely explored without coding expertise. CellCover is available in CellCover R and CellCover Python. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=67 SRC="FIGDIR/small/535943v6_ufig1.gif" ALT="Figure 1"> View larger version (19K): org.highwire.dtl.DTLVardef@893301org.highwire.dtl.DTLVardef@173a6baorg.highwire.dtl.DTLVardef@1c70e7eorg.highwire.dtl.DTLVardef@1888ab3_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Cross-species imputation and comparison of single-cell transcriptomic profiles 96%
- GoM DE: interpreting structure in sequence count data with differential expression analysis allowing for grades of membership 95%
- Integrating temporal single-cell gene expression modalities for trajectory inference and disease prediction 95%
Similar papers in this journal
- Differential detection workflows for multi-sample single-cell RNA-seq data 93%
- Probability of stealth multiplets in sample-multiplexing for droplet-based single-cell analysis 93%
- A Common Methodological Phylogenomics Framework for intra-patient heteroplasmies to infer SARS-CoV-2 sublineages and tumor clones 93%
Similar papers in this journal
- Binomial models uncover biological variation during feature selection of droplet-based single-cell RNA sequencing 96%
- Building, Benchmarking, and Exploring Perturbative Maps of Transcriptional and Morphological Data 95%
- G2S3: a gene graph-based imputation method for single-cell RNA sequencing data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.