Evaluating Sample Augmentation in Microarray Datasets with Generative Models: A Comparative Pipeline and Insights in Tuberculosis.
Gupta, A.; Ahmad, S.; Sune, A.; Gupta, C.; Gulati, H. K.; Kutum, R.; Sethi, T.
Show abstract
High throughput screening technologies have created a fundamental challenge for statistical and machine learning analyses, i.e., the curse of dimensionality. Gene expression data are a quintessential example, high dimensional in variables (Large P) and comparatively much smaller in samples (Small N). However, the large number of variables are not independent. This understanding is reflected in Systems Biology approaches to the transcriptome as a network of coordinated biological functioning or through principal Axes of variation underlying the gene expression. Recent advances in generative deep learning offers a new paradigm to tackle the curse of dimensionality by generating new data from the underlying latent space captured as a deep representation of the observed data. These have led to widespread applications of approaches such as Generative Adversarial Networks (GANs) and Variational Autoencoders (VAEs), especially in domains where millions of data points exist, such as in computer vision and single cell data. Very few studies have focused on generative modeling of bulk transcriptomic data and microarrays, despite being one of the largest types of publicly available biomedical data. Here we review the potential of Generative models in recapitulating and extending biomedical knowledge from microarray data, which may thus limit the potential to yield hundreds of novel biomarkers. Here we review the potential of generative models and conduct a comparative analysis of VAE, GAN and gaussian mixture model (GMM) in a dataset focused on Tuberculosis. We further review whether previously known axes genes can be used as an effective strategy to employ domain knowledge while designing generative models as a means to further reduce biological noise and enhance signals that can be validated by standard enrichment approaches or functional experiments.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- A Convolution Based Computational Approach Towards DNA N6-methyladenine Site Identification and Motif Extraction in Rice Genome 96%
- A novel interpretable deep transfer learning combining diverse learnable parameters for improved T2D prediction based on single-cell gene regulatory networks 95%
- A hybrid CNN-Random Forest algorithm for bacterial spore segmentation and classification in TEM images 94%
Similar papers in this journal
Similar papers in this journal
- Stability of feature selection utilizing Graph Convolutional Neural Network and Layer-wise Relevance Propagation 95%
- AIRBP: Accurate identification of RNA-binding proteins using machine learning techniques 94%
- Graph Neural Network Modelling as a potentially effective Method for predicting and analyzing Procedures based on Patient Diagnoses 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.