A synthetic data generation pipeline to reproducibly mirror high-resolution multi-variable peptidomics and real-patient clinical data
Jaimes Campos, M. A.; Kabic, S.; Latosinska, A.; Anicic, E.; Siwy, J.; Dragusica, V.; Rupprecht, H.; Catanese, L.; Keller, F.; Perco, P.; Gomez - Gomez, E.; Beige, J.; Vlahou, A.; Mischak, H.; Vukelic, D.; Krizan, T.; Frantzi, M.
Show abstract
Generating high quality, real-world clinical and molecular datasets is challenging, costly and time intensive. Consequently, such data should be shared with the scientific community, which however carries the risk of privacy breaches. The latter limitation hinders the scientific communitys ability to freely share and access high resolution and high quality data, which are essential especially in the context of personalised medicine. In this study, we present an algorithm based on Gaussian copulas to generate synthetic data that retain associations within high dimensional (peptidomics) datasets. For this purpose, 3,881 datasets from 10 cohorts were employed, containing clinical, demographic, molecular (> 21,500 peptide) variables, and outcome data for individuals with a kidney or a heart failure event. High dimensional copulas were developed to portray the distribution matrix between the clinical and peptidomics data in the dataset, and based on these distributions, a data matrix of 2,000 synthetic patients was developed. Synthetic data maintained the capacity to reproducibly correlate the peptidomics data with the clinical variables. Consequently, correlation of the rho-values of individual peptides with eGFR between the synthetic and the real-patient datasets was highly similar, both at the single peptide level (rho = 0.885, p < 2.2e-308) and after classification with machine learning models (rhosynthetic = -0.394, p = 5.21e-127; rhoreal = -0.396, p = 4.64e-67). External validation was performed, using independent multi-centric datasets (n = 2,964) of individuals with chronic kidney disease (CKD, defined as eGFR < 60 mL/min/1.73m{superscript 2}) or those with normal kidney function (eGFR > 90 mL/min/1.73m{superscript 2}). Similarly, the association of the rho-values of single peptides with eGFR between the synthetic and the external validation datasets was significantly reproduced (rho = 0.569, p = 1.8e-218). Subsequent development of classifiers by using the synthetic data matrices, resulted in highly predictive values in external real-patient datasets (AUC values of 0.803 and 0.867 for HF and CKD, respectively), demonstrating robustness of the developed method in the generation of synthetic patient data. The proposed pipeline represents a solution for high-dimensional sharing while maintaining patient confidentiality.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Imputation of label-free quantitative mass spectrometry-based proteomics data using self-supervised deep learning 94%
- Mapping specificity, entropy, allosteric changes and substrates in blood proteases by a high-throughput protease screen 94%
- Lipidomic profiling of human serum enables detection of pancreatic cancer 93%
Similar papers in this journal
- Multifaceted proteome analysis at solubility, redox, and expression dimensions for target identification 94%
- ProT-Diff: A Modularized and Efficient Approach to De Novo Generation of Antimicrobial Peptide Sequences through Integration of Protein Language Model and Diffusion Model 92%
- A subset of pro-inflammatory CXCL10+ LILRB2+ macrophages derives from recipient monocytes and drives renal allograft rejection 91%
Similar papers in this journal
Similar papers in this journal
- 1 H-NMR metabolomics-based surrogates to impute common clinical risk factors and endpoints 92%
- Machine learning guided association of adverse drug reactions with in vitro target-based pharmacology 91%
- Integrative deep learning analysis improves colon adenocarcinoma patient stratification at risk for mortality 91%
Similar papers in this journal
- The Interpretable Multimodal Machine Learning (IMML) framework reveals pathological signatures of distal sensorimotor polyneuropathy 95%
- Deep Proteome Profiling of Metabolic Dysfunction-Associated Steatotic Liver Disease 94%
- Proteomic Characterization of Acute Kidney Injury in Patients Hospitalized with SARS-CoV2 Infection 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.