Hidden Knowledge Recovery from GAN-generated Single-cell RNA-seq Data
Zhang, X.; Li, F.; Shah, N. U.
Show abstract
BackgroundMachine learning methods have recently been shown powerful in discovering knowledge from scientific data, offering promising prospects for discovery learning. In the meanwhile, Deep Generative Models like Generative Adversarial Networks (GANs) have excelled in generating synthetic data close to real data. GANs have been extensively employed, primarily motivated by generating synthetic data for privacy preservation, data augmentation, etc. However, certain dimensions of GANs have received limited exploration in current literature. Existing studies predominantly utilize huge datasets, presenting a challenge when dealing with limited, complex datasets. Researchers have high-lighted the ineffectiveness of conventional scores for selecting optimal GANs on limited datasets that exhibit complex high order relationships. Furthermore, current methods evaluate GANs performance by comparing synthetic data to real data without assessing the preservation of high-order relationships. Researchers have advocated for more objective GAN evaluation techniques and emphasized the importance of establishing interpretable connections between GAN latent space variables and meaningful data semantics. ResultsIn this study, we used a custom GAN model to generate quality synthetic data for a very limited, complex biological dataset. We successfully recovered cell-lineage developmental story from synthetic data using the ab-initio knowledge discovery method, we previously developed. Our custom GAN model performed better than state-of-the-art cscGAN model, when evaluated for recovering hidden knowledge from limited, complex dataset. Then we devise a temporal dataset specific quantitative scoring mechanism to successfully reproduce GAN results for human and mouse embryonic datasets. Our Latent Space Interpretation (LSI) scheme was able to identify anomalies. We also found that the latent space in GAN effectively captured the semantic information and may be used to interpolate data when the sampling of real data is sparse. ConclusionIn summary we used a customized GAN model to generate synthetic data for limited, complex dataset and compared the results with state-of-the-art cscGAN model. Cell-lineage developmental story is recovered as hidden knowledge to evaluate GAN for preserving complex high-order relationships. We formulated a quantitative score to successfully reproduce results on human and mouse embryonic datasets. We designed a LSI scheme to identify anomalies and understand the mechanism by which GAN captures important data semantics in its latent space.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Species-Agnostic Transfer Learning for Cross-species Transcriptomics Data Integration without Gene Orthology 95%
- Graph Contrastive Learning as a Versatile Foundation for Advanced scRNA-seq Data Analysis 95%
- CRISPR-DIPOFF: An Interpretable Deep LearningApproach for CRISPR Cas-9 Off-Target Prediction 95%
Similar papers in this journal
- A novel interpretable deep transfer learning combining diverse learnable parameters for improved T2D prediction based on single-cell gene regulatory networks 96%
- Unsupervised generative and graph representation learning for modelling cell differentiation 95%
- Predicting cell types with supervised contrastive learning on cells and their types 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.