A deep generative model for capturing cell to phenotype relationships
Weinberger, E.; Yu, P.; Lee, S.-I.
Show abstract
Single-cell omics has proven to be a powerful instrument for exploring cellular diversity. With advances in sequencing protocols, single-cell studies are now routinely collected from large-scale donor cohorts consisting of samples from hundreds of donors with the goal of uncovering the molecular bases of higher-level donor phenotypes of interest. For example, to better understand the mechanisms behind Alzheimers disease, recent studies with up to hundreds of samples have investigated the relationships between single-cell omics measurements and donors neuropathological phenotypes (e.g. Braak staging) [4, 9, 3, 10]. In order to ensure the robustness of such findings, it may be desirable to aggregate data from multiple distinct donor cohorts. Unfortunately, doing so is not always straightforward, as different cohorts may be equipped with different sets of phenotype labels. Continuing the previous Alzheimers example, recent AD study cohorts have provided various subsets of neuropathological phenotypes, cognitive testing results, and APOE genotype. Thus, it is desirable to be able to infer any missing phenotype labels such that all available cell-level data in the study of a given phenotype of interest could be used. Moreover, beyond simply imputing missing phenotype information, it is often of interest to understand which groups of cells and/or molecular features may be most predictive of a given phenotype of interest. As such, there is a pressing need for computational methods that can connect cell-level measurements with donor-level labels. However, accomplishing this task is not straightforward. While a rich literature exists on learning meaningful low-dimensional representations of cells [7, 8, 1, 2] and for inferring corresponding cell-level labels (e.g. cell type) [11], the donor level prediction task introduces substantial additional complexity. For example, different numbers of cells may be recovered from each donor, and thus our prediction model must be able to handle arbitrary numbers of samples as input. Moreover, ideally our model would not a priori require any additional prior knowledge beyond our cell-level measurements, such as the importance of different cell types for a given prediction task. To resolve these issues, here we propose milVI (multiple instance learning variational inference), a deep generative modeling framework that explicitly accounts for donor-level phenotypes and enables inference of missing phenotype labels post-training. In order to handle varying numbers of cells per donor when inferring phenotype labels, milVI leverages recent advances in multiple instance learning. We validated milVI by applying to impute held-out Braak staging information from an Alzheimers disease study cohort from Mathys et al. [9], and we found that our method achieved lower error on this task compared to naive imputation methods.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deep feature extraction of single-cell transcriptomes by generative adversarial network 95%
- TUGDA: Task uncertainty guided domain adaptation for robust generalization of cancer drug response prediction from in vitro to in vivo settings 94%
- ARTEMIS integrates autoencoders and schrodinger bridges to predict continuous dynamics of gene expression, cell population and perturbation from time-series single-cell data 94%
Similar papers in this journal
- Hypergraph factorisation for multi-tissue gene expression imputation 95%
- Simultaneous dimensionality reduction and integration for single-cell ATAC-seq data using deep learning 95%
- Inferring spatial single-cell-level interactions through interpreting cell state and niche correlations learned by self-supervised graph transformer 94%
Similar papers in this journal
- Learning interpretable cellular and gene signature embeddings from single-cell transcriptomic data 95%
- scMODAL: A general deep learning framework for comprehensive single-cell multi-omics data alignment with feature links 94%
- scDREAMER: atlas-level integration of single-cell datasets using deep generative model paired with adversarial classifier 94%
Similar papers in this journal
Similar papers in this journal
- Hierarchical confounder discovery in the experiment-machine learning cycle 93%
- Single-Cell Multi-Modal GAN (scMMGAN) reveals spatial patterns in single-cell data from triple negative breast cancer 93%
- Generating hard-to-obtain information from easy-to-obtain information: applications in drug discovery and clinical inference 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.