Towards Useful and Private Synthetic Omics: Community Benchmarking of Generative Models for Transcriptomics Data
Öztürk, H.; Afonja, T.; Jälkö, J.; Binkyte, R.; Rodriguez-Mier, P.; Lobentanzer, S.; Wicks, A.; Kreuer, J.; Ouaari, S.; Pfeifer, N.; Menzies, S.; Pentyala, S.; Filienko, D.; Golob, S.; McKeever, P.; Banerjee, J.; Foschini, L.; De Cock, M.; Saez-Rodriguez, J.; Fritz, M.; Stegle, O.; Honkela, A.
Show abstract
BackgroundThe synthesis of anonymized data derived from real-world cohorts offers a promising strategy for regulatory-compliant and privacy-preserving biological data sharing, potentially facilitating model development that can improve predictive performance. However, the extent to which generative models can preserve biological signals while remaining resilient to adversarial privacy attacks in high-dimensional omics contexts remains underexplored. To address this gap, the CAMDA 2025 Health Privacy Challenge launched a community-driven effort to systematically benchmark synthetic and privacy-preserving data generation for bulk RNA-seq cohorts. ResultsBuilding on this initiative, we systematically benchmarked 11 generative methods across two cancer cohorts ([~]1,000 and [~]5,000 patients) over 978 landmark genes. Methods were evaluated across complementary axes of distributional fidelity, downstream utility, biological plausibility and empirical privacy risk, with emphasis on trade-offs between vulnerability to membership inference attacks (MIA) and other evaluation dimensions. Expressive deep generative models achieved strong predictive utility and differential expression recovery, but were often more vulnerable to membership inference risk. Differentially private methods improved resistance to attacks at the cost of reduced utility, while simpler statistical approaches offered competitive utility with moderate privacy risk and fast training. ConclusionsSynthetic bulk RNA-seq quality is inherently multi-dimensional and shaped by trade-offs between utility, biological preservation and privacy. Our results indicate that differences in model architecture drive distinct trade-offs across these axes, suggesting that model choice should align with dataset characteristics, intended downstream use and privacy requirements. Privacy risk should also be assessed using multiple complementary attack methods and, where possible, formal differential privacy protection.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- GRouNdGAN: GRN-guided simulation of single-cell RNA-seq data using causal generative adversarial networks 96%
- scPRINT: pre-training on 50 million cells allows robust gene network predictions 96%
- scConfluence : single-cell diagonal integration with regularized Inverse Optimal Transport on weakly connected features 95%
Similar papers in this journal
- JIND: Joint Integration and Discrimination for Automated Single-Cell Annotation 96%
- EvoAug-TF: Extending evolution-inspired data augmentations for genomic deep learning to TensorFlow 95%
- TUGDA: Task uncertainty guided domain adaptation for robust generalization of cancer drug response prediction from in vitro to in vivo settings 94%
Similar papers in this journal
- Spatial gene expression at single-cell resolution from histology using deep learning with GHIST 95%
- scGPT: Towards Building a Foundation Model for Single-Cell Multi-omics Using Generative AI 95%
- Deep learning-based predictions of gene perturbation effects do not yet outperform simple linear baselines 95%
Similar papers in this journal
- Deep convolutional and conditional neural networks for large-scale genomic data generation 95%
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 95%
- Accessible, Reproducible, and Scalable Machine Learning for Biomedicine 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.