Benchmarking Variational AutoEncoders on cancer transcriptomics data
elTager, M.; Abdelaal, T.; Charrout, M.; Mahfouz, A.; Reinders, M.; Makrodimitris, S.
Show abstract
Deep generative models, such as variational autoencoders (VAE), have gained increasing attention in computational biology due to their ability to capture complex data manifolds which subsequently can be used to achieve better performance in downstream tasks, such as cancer type prediction or subtyping of cancer. However, these models are difficult to train due to the large number of hyperparameters that need to be tuned. To get a better understanding of the importance of the different hyperparameters, we examined six different VAE models when trained on TCGA transcriptomics data and evaluated on the downstream task of cluster agreement with cancer subtypes. We studied the effect of the latent space dimensionality, learning rate, optimizer and initialization on the quality of subsequent clustering of the TCGA samples. We found {beta}-TCVAE and DIP-VAE to have a good performance, on average, despite being more sensitive to hyperparameters selection. Based on these experiments, we derived recommendations for selecting the different hyperparameters settings. In addition, we examined whether the learned latent spaces capture biologically relevant information. Hereto, we correlated the different representations with various data characteristics such as age, days to metastasis, immune infiltration, and mutation signatures. We found that for all models the latent factors, in general, do not uniquely correlate with one of the data characteristics even for models specifically designed for disentanglement.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Species-Agnostic Transfer Learning for Cross-species Transcriptomics Data Integration without Gene Orthology 97%
- Graph Contrastive Learning as a Versatile Foundation for Advanced scRNA-seq Data Analysis 96%
- Synthetic observations from deep generative models and binary omics data with limited sample size 96%
Similar papers in this journal
Similar papers in this journal
- Learning universal knowledge graph embedding for predicting biomedical pairwise interactions 95%
- iDRKAN: Interpretable miRNA-Disease Association Prediction Based on Dual-Graph Representation Learning and Kolmogorov-Arnold Network 94%
- GAT-HiC: Efficient Reconstruction of 3D Chromosome Structure via Residual Graph Attention Neural Networks 94%
Similar papers in this journal
- Generative AI Mitigates Representation Bias Using Synthetic Health Data 96%
- A variational autoencoder trained with priors from canonical pathways increases the interpretability of transcriptome data 96%
- Learning massive interpretable gene regulatory networks of the human brain by merging Bayesian Networks 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.