Sequential compression of gene expression across dimensionalities and methods reveals no single best method or dimensionality
Way, G. P.; Zietz, M.; Rubinetti, V.; Himmelstein, D. S.; Greene, C. S.
Show abstract
BackgroundUnsupervised compression algorithms applied to gene expression data extract latent, or hidden, signals representing technical and biological sources of variation. However, these algorithms require a user to select a biologically-appropriate latent dimensionality. In practice, most researchers select a single algorithm and latent dimensionality. We sought to determine the extent by which using multiple dimensionalities across ensemble compression models improves biological representations.\n\nResultsWe compressed gene expression data from three large datasets consisting of adult normal tissue, adult cancer tissue, and pediatric cancer tissue. We compressed these data into many latent dimensionalities ranging from 2 to 200. We observed various tradeoffs across latent dimensionalities and compression models. For example, we observed high model stability between principal components analysis (PCA), independent components analysis (ICA), and non-negative matrix factorization (NMF). We identified more unique biological signatures in ensembles of denoising autoencoder (DAE) and variational autoencoder (VAE) models in intermediate latent dimensionalities. However, we captured the most pathway-associated features using all compressed features across algorithms and dimensionalities. Optimized at different latent dimensionalities, compression models detect generalizable gene expression signatures representing sex, neuroblastoma MYCN amplification, and cell types. In two supervised machine learning tasks, compressed features optimized predictions at different latent dimensionalities.\n\nConclusionsThere is no single best latent dimensionality or compression algorithm for analyzing gene expression data. Instead, using feature ensembles from different compression models across latent space dimensionalities optimizes biological representations.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- FastCCC: A permutation-free framework for scalable, robust, and reference-based cell-cell communication analysis in single cell transcriptomics studies 97%
- multiDGD: A versatile deep generative model for multi-omics data 96%
- Atlas-scale single-cell multi-sample multi-condition data integration using scMerge2 96%
Similar papers in this journal
- Building, Benchmarking, and Exploring Perturbative Maps of Transcriptional and Morphological Data 96%
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 96%
- Optimal tuning of weighted kNN- and diffusion-based methods for denoising single cell genomics data 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.