Dirichlet process mixture models to estimate outcomes for individuals with missing predictor data: application to predict optimal type 2 diabetes therapy in electronic health record data
Cardoso, P.; Dennis, J. M.; Bowden, J.; Shields, B.; McKinley, T.; MASTERMIND Consortium,
Show abstract
BackgroundMissing data is a common problem in regression modelling. Much of the literature focuses on handling missing outcome variables, but there are also challenges when dealing with missing predictor information, particularly when trying to build prediction models for use in practice. MethodsWe develop a flexible Bayesian approach for handling missing predictor information in regression models. For prediction this provides practitioners with full posterior predictive distributions for both the missing predictor information and the outcome variable, conditional on the observed predictors. We apply our approach to a previously proposed treatment selection model for type 2 diabetes second-line therapies. Our approach combines a regression model and a Dirichlet process mixture model (DPMM), where the former defines the treatment selection model and the latter provides a flexible way to model the joint distribution of the predictors. ResultsWe show that under missing-completely-at-random (MCAR) and missing-at-random (MAR) assumptions (with respect to the missing predictors), the DPMM can model complex relationships between predictor variables, and predict missing values conditionally on existing information. We also demonstrate that in the presence of multiple missing predictors, the DPMM model can be used to explore which variable(s), if collected, could provide the most additional information about the likely outcome. ConclusionsOur approach can provide practitioners with supplementary information to aid treatment selection decisions in the presence of missing data, and can be readily extended to other types of response model. Key MessagesO_LIMissing predictor variables present a significant challenge when building and implementing prediction models in clinical practice. C_LIO_LIRemoving individuals with missing information and performing a complete case analysis can lead to imprecision and bias. Multiple imputation approaches typically translate uncertainty through prediction model parameter standard errors, as opposed to a consistent joint probability model. C_LIO_LIAlternatively, a Bayesian approach using Dirichlet process mixture models (DPMMs) offers a flexible way to model complex joint distributions of predictor variables, which can be used to estimate posterior (predictive) distributions for the missing predictors, conditional on the observed predictors. C_LIO_LIUsing a DPMM, in this way allows uncertainties around missing predictor data to be propagated through to a prediction model of interest using a Bayesian hierarchical framework. This allows prediction models to be developed using datasets with incomplete predictor information (assuming missing-completely-at-random/missing-at-random). Furthermore, predictions can be made on new individuals even if they have incomplete predictor information (under the same assumptions). C_LIO_LIThis approach provides full posterior predictive probability distributions for both missing predictor variables and the outcome variable, allowing a wide range of probabilistic models outputs to be derived to support clinical decision making. C_LI
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Stability-Enhanced Lasso Approach for Covariate Selection in Non-Linear Mixed Effect Model 95%
- Penalized reduced rank regression for multi-outcome survival data supports a common metabolic risk score for age-related diseases 94%
- Using generalized additive models to analyze biomedical non-linear longitudinal data 93%
Similar papers in this journal
- Bayesian Structural Time Series for Biomedical Sensor Data: A Flexible Modeling Framework for Evaluating Interventions 94%
- A regularized functional regression model enabling transcriptome-wide dosage-dependent association study of cancer drug response 93%
- Model guided trait-specific co-expression network estimation as a new perspective for identifying molecular interactions and pathways 93%
Similar papers in this journal
- External control arm analysis: an evaluation of propensity score approaches, G-computation, and doubly debiased machine learning 94%
- Prediction-powered Inference for Clinical Trials 94%
- Comparison of Bayesian networks, G-estimation and linear models to estimate causal treatment effects in aggregated N-of-1 trials 94%
Similar papers in this journal
- Optimization of nutritional strategies using a mechanistic computational model in prediabetes: Application to the J-DOIT1 study data 93%
- Predicting diabetes second-line therapy initiation in the Australian population via timespan-guided neural attention network 92%
- Applying Historical Data in a Nonlinear Mixed-Effects Model Can Reduce the Number of Control Rats Required for Calculation of the Relative Potency of Insulin Analogues 92%
Similar papers in this journal
- Individual Reference Intervals for Personalized Interpretation of Clinical and Metabolomics Measurements 95%
- A methodology of phenotyping ICU patients from EHR data: high-fidelity, personalized, and interpretable phenotypes estimation 94%
- Automated Interpretable Discovery of Heterogeneous Treatment Effectiveness: A Covid-19 Case Study 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.