Validation of Generative AI Techniques for Synthetic Data Generation in Multiple Sclerosis Research: A Comparison with Real-World Evidence from the Italian MS Registry
Iaffaldano, P.; D'Amico, S.; Lucisano, G.; Copetti, M.; Guerra, T.; Rocca, M. A.; Patti, F.; De Luca, G.; Ferraro, D.; Totaro, R.; Brescia Morra, V.; Salemi, G.; Portaccio, E.; Foschi, M.; Inglese, M.; Coniglio, M. G.; Chisari, C. G.; Caputo, F.; Paolicelli, D.; Battaglia, M. A.; Della Porta, M. G.; Savevski, V.; Delleani, M.; Colella, F. E.; Sauta, E.; Amato, M. P.; Filippi, M.; Trojano, M.
Show abstract
ImportanceLarge multiple sclerosis (MS) registries provide crucial real-world evidence but often suffer from missing data, inconsistencies, and privacy limitations that restrict data sharing. The use of generative AI to create synthetic data (SD) is an emerging strategy to enhance real-world evidence research potentially overcoming these challenges. ObjectiveTo evaluate the validity of AI-generated synthetic data (SD) in replicating real data collected in the Italian MS and Related Disorders Register (RISM), and to compare the risk of progression independent of relapse activity (PIRA) between early intensive treatment (EIT) versus escalation treatment strategy (ESC) in both real and synthetic MS cohorts. Design, Setting, and ParticipantsThis validation study analyzed data from RISM. AI-based generative models were trained on a sub-cohort of 1,666 patients with tabularized MRI data to generate a synthetic dataset of 4,878 patients. SD was evaluated using the Synthetic vAlidation FramEwork powered by Train (SAFE), assessing fidelity, utility, and privacy. Clinical Synthetic Fidelity (CSF) and Nearest Neighbor Distance Ratio (NNDR) were used for statistical and privacy validation. Treatment outcome comparisons between EIT and ESC strategies were conducted for clinical validation using both real and synthetic datasets, focusing on the risk of PIRA. ExposuresInitial disease-modifying therapy strategy, categorized as EIT versus ESC. Main Outcomes and MeasuresPrimary outcome was the occurrence of PIRA, defined as confirmed disability accrual independent of relapses. Validation metrics included Clinical Synthetic Fidelity (CSF [≥]90 optimal) and Nearest Neighbor Distance Ratio (NNDR, range 0.60-0.85 for privacy). ResultsThe synthetic dataset demonstrated high fidelity (CSF=97%) and privacy preservation (NNDR=0.61). Treatment effect estimates for ESCs vs EIT were consistent across real and synthetic datasets, with largely comparable trends, with increased statistical significance in SD. Cox proportional hazards models confirmed the robustness of synthetic data in estimating the risk of the first PIRA event. Conclusions and RelevanceAI-generated synthetic data reliably replicated treatment effect outcomes from real-world RISM data, overcoming missing data and providing a privacy-preserving alternative for data sharing and clinical research. Key pointsO_ST_ABSQuestionC_ST_ABSCan Artificial Intelligence (AI)-generated synthetic data (SD) reliably replicate multiple sclerosis (MS) registry data and provide robust insights into progression independent of relapse activity (PIRA) phenomena under different treatment strategies? FindingsIn a cohort of 4,878 relapsing-onset MS patients from the Italian MS Register, AI-generated SD achieved high fidelity (CSF = 97%), and reproduced treatment effect outcomes. Both real and synthetic cohorts consistently showed that early intensive therapy reduced the risk of PIRA compared with an escalation strategy. MeaningSD can complement and enhance registry-based research by addressing missing data and supporting reproducible analyses in MS.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Creating an automated tool for a consistent and repeatable evaluation of disability progression in clinical studies for Multiple Sclerosis 94%
- Tissue damage detected by quantitative gradient echo MRI correlates with clinical progression in non-relapsing progressive MS 93%
- Disability patterns in multiple sclerosis: a meta-analysis on PIRA and RAW in the real world context 92%
Similar papers in this journal
- The Interpretable Multimodal Machine Learning (IMML) framework reveals pathological signatures of distal sensorimotor polyneuropathy 90%
- LUNAR: A Deep Learning Model to Predict Glioma Recurrence Using Integrated Genomic and Clinical Data 90%
- Pretrained Patient Trajectories for Adverse Drug Event Prediction Using Common Data Model-based Electronic Health Records 89%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.