Toward Reliable Synthetic Omics: Statistical Distances for Generative Models Evaluation
Marchesi, R.; Lazzaro, N.; Leonardi, G.; Rignanese, F.; Bovo, S.; Chierici, M.; Jurman, G.
Show abstract
BackgroundSynthetic data generation is emerging as an approach to overcome the limitations of real-world data scarcity in omics studies, especially in precision medicine and oncology. Omics datasets, with their high dimensionality and relatively small sample sizes, often lead to overfitting, especially in deep learning models. Generative models offer a promising way to generate realistic synthetic data preserving the original data distribution. However, there is still no objective consensus on how to evaluate their performance. In this study, we set out to validate generative networks for transcriptomics data generation by using statistical distances as robust evaluation metrics. ResultsWe observe that statistical distances enable simultaneous evaluation of global and local data fidelity of generated synthetic data. Because these distances satisfy the properties of true metrics, they also enable formal hypothesis testing to assess whether generative models have in fact converged or are merely approaching the reference distribution. Crucially, optimizing for these distances was found to implicitly select models maximizing other widely used metrics of generative performance, providing evidence of their broad applicability. Overall, our findings indicate that the adoption of these metrics can play a key role in guiding the development of generative models across a wide range of domains.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- ACTIVA: realistic single-cell RNA-seq generation with automatic cell-type identification using introspective variational autoencoders 97%
- scVAE: Variational auto-encoders for single-cell gene expression data 97%
- Adversarial Deconfounding Autoencoder for Learning Robust Gene Expression Embeddings 96%
Similar papers in this journal
- Synthetic observations from deep generative models and binary omics data with limited sample size 97%
- Evaluation of out-of-distribution detection methods for data shifts in single-cell transcriptomics 96%
- Evaluating discrepancies in dimensionality reduction for time-series single-cell RNA-sequencing data 95%
Similar papers in this journal
- Tree-informed Bayesian multi-source domain adaptation: cross-population probabilistic cause-of-death assignment using verbal autopsy 94%
- Survival Analysis on Rare Events Using Group-Regularized Multi-Response Cox Regression 94%
- The winner's curse under dependence: repairing empirical Bayes using convoluted densities 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.