Latent Feature Representations for Human Gene Expression Data Improve Phenotypic Predictions
Pantazis, Y.; Tselas, C.; Lakiotaki, K.; Lagani, V.; Tsamardinos, I.
Show abstract
High-throughput technologies such as microarrays and RNA-sequencing (RNA-seq) allow to precisely quantify transcriptomic profiles, generating datasets that are inevitably high-dimensional. In this work, we investigate whether the whole human transcriptome can be represented in a compressed, low dimensional latent space without loosing relevant information. We thus constructed low-dimensional latent feature spaces of the human genome, by utilizing three dimensionality reduction approaches and a diverse set of curated datasets. We applied standard Principal Component Analysis (PCA), kernel PCA and Autoencoder Neural Networks on 1360 datasets from four different measurement technologies. The latent feature spaces are tested for their ability to (a) reconstruct the original data and (b) improve predictive performance on validation datasets not used during the creation of the feature space. While linear techniques show better reconstruction performance, nonlinear approaches, particularly, neural-based models seem to be able to capture non-additive interaction effects, and thus enjoy stronger predictive capabilities. Our results show that low dimensional representations of the human transcriptome can be achieved by integrating hundreds of datasets, despite the limited sample size of each dataset and the biological / technological heterogeneity across studies. The created space is two to three orders of magnitude smaller compared to the raw data, offering the ability of capturing a large portion of the original data variability and eventually reducing computational time for downstream analyses.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- A variational autoencoder trained with priors from canonical pathways increases the interpretability of transcriptome data 97%
- A Generalized Higher-order Correlation Analysis Framework for Multi-Omics Network Inference 96%
- Detection of genes with differential expression dispersion unravels the role of autophagy in cancer progression 96%
Similar papers in this journal
- DeepInsight-3D for precision oncology: an improved anti-cancer drug response prediction from high-dimensional multi-omics data with convolutional neural networks 97%
- Tensor decomposition- and principal component analysis-based unsupervised feature extraction to select more reasonable differentially expressed genes: Optimization of standard deviation versus state-of-art methods 96%
- Finding disease modules for cancer and COVID-19 in gene co-expression networks with the Core&Peel method 96%
Similar papers in this journal
- Molecular Group and Correlation Guided Structural Learning for Multi-Phenotype Prediction 97%
- Species-Agnostic Transfer Learning for Cross-species Transcriptomics Data Integration without Gene Orthology 96%
- Assessing Random Forest self-reproducibility for optimal short biomarker signature discovery 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.