Generate Synthetic Data in R for a Hypothetical Alzheimer's Disease Trial
Handels, R.; Jonsson, L.; Raket, L. L.; Alzheimer's Disease Neuroimaging Initiative,
Show abstract
INTRODUCTIONRepresentative data of recent Alzheimers Disease (AD) trials are difficult to obtain. We aimed to generate a synthetic version of an original real-world observational dataset, subsequently apply a plausible AD treatment effect, and make our method open-source available. METHODSSynthetic data was generated in the following steps: (1) Obtain real-world data from the ADNI study on demographic (age, sex, education), clinical (cognition: MMSE and ADAS; function: FAQ; composite cognition/function: CDR, ADCOMS) and biological (genetics: APOE4; cerebrospinal fluid: ABeta, Tau; imaging: PET-SUVR-centiloid) outcomes at baseline, 6, 12 and/or 18-month follow-up (35 variables), with missing data multiple-imputed to obtain 10 sets of 537 individuals. (2) Estimate (theoretical) minimum and maximum (all continuous variables) and proportions (all categorical variables). (3) Rescale to 0-1 range (continuous). (4) Estimate beta distribution shape parameters (method of moments; continuous). (5) Transform to cumulative probability distribution function (using shape parameters; continuous) and to cumulative probability (categorical). (6) Transform to a normal distribution. (7) Estimate variance-covariance matrix. (8) Generate random correlated normal data using Cholesky decomposition of variance-covariance. (9) Transform to cumulative probability distribution function. (10) Transform to beta distribution (using shape parameters; continuous). (11) Rescale to original range. (12) Keep half as control arm, and half as intervention arm, and estimate change from baseline. (13) Multiply intervention change from baseline with self-defined hypothetical relative treatment effect. We assumed correlations on normalized scale were similar to correlations on original scale. R code is available on github: https://github.com/ronhandels/synthetic-correlated-data. RESULTSThe synthetic distribution and mean over time showed large similarity to the original data (visually assessed). The absolute difference in pairwise correlations between original and synthetic data median was 0.02 (95th percentile=0.11, max=0.18). CONCLUSIONWe judged our method sufficiently valid to generate synthetic correlated plausible hypothetical trial results.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- c-Triadem: A constrained, explainable deep learning model to identify novel biomarkers in Alzheimer’s disease 94%
- Random forest model for feature-based Alzheimer's disease conversion prediction from early mild cognitive impairment subjects 94%
- Interpretable multivariate survival models: Improving predictions for conversion from mild cognitive impairment to Alzheimers disease (AD) via data fusion and machine learning 93%
Similar papers in this journal
- Returning Research Results That Indicate Risk of Alzheimer Disease Dementia to Healthy Participants in Longitudinal Studies (WeSHARE) 93%
- The New Therapeutics in Alzheimer’s Disease Longitudinal Cohort study (NTAD): study protocol 93%
- Optimising activity and diet compositions for dementia prevention: Protocol for the ACTIVate prospective longitudinal cohort study 92%
Similar papers in this journal
- Screening for early-stage Alzheimer's disease using optimized feature sets and machine learning 94%
- Exploring the Correlation between the Cognitive Benefits of Drug Combinations in a Clinical Alzheimer Disease Database and the Efficacies of the Same Drug Combinations Predicted from a Computational Model 94%
- Topographical overlapping of the Aβ and Tau pathologies in the Default mode networks predicts Alzheimer’s Disease with higher specificity 93%
Similar papers in this journal
- Unraveling Alzheimer's Disease: Investigating Dynamic Functional Connectivity in the Default Mode Network through DCC-GARCH Modeling 91%
- QRATER: a collaborative and centralized imaging quality control web-based application. 91%
- Visual QC Protocol for FreeSurfer Cortical Parcellations from Anatomical MRI 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.