Back

TabSyM: A Generative Pipeline for Small Multi-Cohort Omics Tabular Data

Yu, N.; Wang, Y.; Olsen, L. K.; Zhang, B.; Zhang, h.; Liu, Z.

2025-07-18 bioinformatics
10.1101/2025.07.14.664738 bioRxiv
Show abstract

Machine learning applications in biomedicine such as omics data analysis are frequently hindered by datasets that are small, high-dimensional, and affected by batch effects across different patient cohorts. To address these challenges, we introduce TabSyM, a modular generative pipeline that synthesizes high-quality, task-relevant data to improve predictive modeling. TabSyM integrates three key stages: it extends a diffusion-based model (TabDDPM) to generate new omics data, employs a novel task-aware sampling mechanism guided by Bayesian optimization to select the most informative synthetic samples, and uses a Multi-Domain Adversarial Network (MDAN) to align data distributions for cross-cohort generalization. We validated our pipeline on a challenging, real-world task of predicting 3-year survival in gastric cancer patients from high-dimensional scRNA-seq data across five cohorts. The full TabSyM pipeline achieved a 30.2% AUROC improvement over the best tree-based models and an 11.5% AUROC gain over leading automated machine learning frameworks. Furthermore, the generative and sampling components are model-agnostic and can substantially boost the performance of classical models like XGBoost independently. These results establish that combining generative modeling with task-aware sampling and domain adaptation provides a robust and effective strategy for overcoming critical data limitations in biomedical tabular data analysis.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.