Multiple imputation as an alternative method to simulate covariate distributions: comparison with joint multivariate normal and bootstrap techniques
Smania, G.; Jonsson, N. E.
Show abstract
Clinical trial simulation (CTS) is a valuable tool in drug development. To obtain realistic scenarios, the subjects included in the CTS must be representative of the target population. Common ways of generating virtual subjects are based upon bootstrap (BS) procedures or multivariate normal distributions (MVND). Here, we investigated the performance of an alternative method based on conditional distributions (CD). Covariates data from a hypertension drug development program were used. The methods were evaluated based on the original dataset (internal evaluation) and on their ability to reproduce an older, unobserved population (extrapolation). Similar results were obtained in the internal evaluation for summary statistics, yet BS was able to preserve the correlation structure of the empirical distribution, which was not adequately reproduced by MVND; CD was in between BS and MVND. BS does not allow to extrapolate to an unobserved population. When the dataset used to inform the extrapolation was well approximated by a MVND, the results from CD and MVND were comparable. However, improved extrapolation performance was observed for CD when deviations from normality assumptions occurred. If CTS is used to simulate within the observed distribution, BS is the preferred method. When extrapolating to new populations, a parametric method like CD/MVND is needed. In case the empirical multivariate distribution is characterized by linearly related covariates and unimodal marginal distributions, MVND can be used because of the simpler statistical framework and well-established use; however, if uncertainty about the MVND assumptions exists, CD will increase the confidence in the simulations compared to MVND.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A regularized functional regression model enabling transcriptome-wide dosage-dependent association study of cancer drug response 93%
- A time-series analysis of blood-based biomarkers within a 25-year longitudinal dolphin cohort. 93%
- Model guided trait-specific co-expression network estimation as a new perspective for identifying molecular interactions and pathways 93%
Similar papers in this journal
- Latent class regression improves the predictive acuity and clinical utility of survival prognostication amongst chronic heart failure patients. 93%
- A machine learning approach to identifying important features for achieving step thresholds in individuals with chronic stroke 93%
- Optimization of nutritional strategies using a mechanistic computational model in prediabetes: Application to the J-DOIT1 study data 92%
Similar papers in this journal
Similar papers in this journal
- A Stability-Enhanced Lasso Approach for Covariate Selection in Non-Linear Mixed Effect Model 95%
- Using generalized additive models to analyze biomedical non-linear longitudinal data 94%
- Penalized reduced rank regression for multi-outcome survival data supports a common metabolic risk score for age-related diseases 94%
Similar papers in this journal
- Can machine learning improve risk prediction of incident hypertension? An internal method comparison and external validation of the Framingham risk model using HUNT Study data 93%
- Machine learning for classifying chronic kidney disease and predicting creatinine levels using at-home measurements 92%
- Predicting bloodstream infection outcome using machine learning 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.