Quantitative Comparison of Structural and Distributional Properties of Synthetic Tabular Data in Parkinson's Disease
Hu, B.; Warsif, S.; Raza, F.; Patel, d.; Chomiak, T.
Show abstract
BackgroundParkinsons disease (PD) research relies heavily on patient data, but access is often limited by privacy concerns, data scarcity, and collection costs. Synthetic data generation offers a potential solution, but its utility hinges on rigorously evaluated fidelity to real-world data. This study quantitatively assesses the structural and distributional fidelity of synthetic tabular data designed to represent PD patients. MethodsWe compared a synthetically generated dataset (N=500 hypothetical entries) against an anonymized real-world dataset (N=57 PD patients) containing demographics, clinical scores (UPDRS, MoCA), and mobility data (6MWT-related variables). The evaluation focused on three key quantitative metrics: (1) Column Correlation Stability, measured by the average absolute difference between Pearson correlation matrices, assessed overall and for clinically relevant variable subgroups (6MWT, UPDRS, MoCA); (2) Principal Component Analysis (PCA), evaluating the variance captured by the top principal components in both datasets; and (3) Jensen-Shannon Distance (JSD), quantifying the distributional similarity between real and synthetic variables across different groups. ResultsThe overall average absolute correlation difference between the real and synthetic datasets was 0.049, indicating moderate preservation of pairwise variable relationships globally. However, stability varied across subgroups, with the 6MWT group showing higher fidelity (difference [~]0.044) compared to the UPDRS ([~]0.080) and MoCA (0.081) groups. PCA revealed that the first two principal components captured 21.36% and 16.36% of the variance, respectively, with visual analysis showing partial overlap between real and synthetic data clusters. Average JSD values indicated moderate distributional similarity overall, with the MoCA group exhibiting the highest fidelity (JSD = 0.0573), while Demographics (0.1167), Clinical (0.1256), and 6MWT (0.1175) groups showed lower distributional similarity. ConclusionSynthetic data generation techniques can replicate univariate distributional properties of PD patient data with moderate success, particularly for certain variable types like cognitive assessments (MoCA). However, accurately capturing the complex multivariate correlation structures, crucial for understanding symptom interactions and building predictive models, remains a significant challenge, especially within specific clinical domains like UPDRS. While synthetic data holds promise for addressing data access issues in PD research, particularly for tasks less sensitive to correlation structure, its application requires careful, context-specific validation. Further development is needed to enhance the structural fidelity of synthetic tabular data for high-stakes, multivariate clinical research applications.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Autocorrelation-based method to identify disordered rhythm in Parkinsons disease tasks: a novel approach applicable to multimodal devices 93%
- Connecting real-world digital mobility assessment to clinical outcomes for regulatory and clinical endorsement – the Mobilise-D study protocol 93%
- A machine learning approach to identifying important features for achieving step thresholds in individuals with chronic stroke 93%
Similar papers in this journal
- Quantifying Device Type and Handedness Biases in a Remote Parkinson’s Disease AI-Powered Assessment 96%
- Crowdsourcing digital health measures to predict Parkinson's disease severity: the Parkinson's Disease Digital Biomarker DREAM Challenge 94%
- A Machine-Learning Based Objective Measure for ALS Disease Severity 92%
Similar papers in this journal
- Characterizing subgroup performance of probabilistic phenotype algorithms within older adults: A case study for dementia, mild cognitive impairment, and Alzheimer’s and Parkinson’s diseases 93%
- Trajectories: a framework for detecting temporal clinical event sequences from health data standardized to the OMOP Common Data Model 92%
- Modeling physician variability to prioritize relevant medical record information 91%
Similar papers in this journal
- An explainable spatial-temporal graphical convolutional network to score freezing of gait in parkinsonian patients 96%
- Motor signatures in digitized cognitive and memory tests enhances characterization of Parkinson’s disease 94%
- Development of a Tremor Detection Algorithm for use in an Academic Movement Disorders Center 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.