Comparison of Bayesian approaches for developing prediction models in rare disease: application to the identification of patients with Maturity-Onset Diabetes of the Young
Cardoso, P.; McDonald, T. J.; Patel, K. A.; Pearson, E. R.; Hattersley, A. T.; Shields, B. M.; McKinley, T. J.
Show abstract
BackgroundClinical prediction models can help identify high-risk patients and facilitate timely interventions. However, developing such models for rare diseases presents challenges due to the scarcity of affected patients for developing and calibrating models. Methods that pool information from multiple sources can help with these challenges. MethodsWe compared three approaches for developing clinical prediction models for population-screening based on an example of discriminating a rare form of diabetes (Maturity-Onset Diabetes of the Young - MODY) in insulin-treated patients from the more common Type 1 diabetes (T1D). Two datasets were used: a case-control dataset (278 T1D, 177 MODY) and a population-representative dataset (1418 patients, 96 MODY tested with biomarker testing, 7 MODY positive). To build a population-level prediction model, we compared three methods for recalibrating models developed in case-control data. These were prevalence adjustment ("offset"), shrinkage recalibration in the population-level dataset ("recalibration"), and a refitting of the model to the population-level dataset ("re-estimation"). We then developed a Bayesian hierarchical mixture model combining shrinkage recalibration with additional informative biomarker information only available in the population-representative dataset. We developed prior information from the literature and other data sources to deal with missing biomarker and outcome information and to ensure the clinical validity of predictions for certain biomarker combinations. ResultsThe offset, re-estimation, and recalibration methods showed good calibration in the population-representative dataset. The offset and recalibration methods displayed the lowest predictive uncertainty due to borrowing information from the fitted case-control model. We demonstrate the potential of a mixture model for incorporating informative biomarkers, which significantly enhanced the models predictive accuracy, reduced uncertainty, and showed higher stability in all ranges of predictive outcome probabilities. ConclusionWe have compared several approaches that could be used to develop prediction models for rare diseases. Our findings highlight the recalibration mixture model as the optimal strategy if a population-level dataset is available. This approach offers the flexibility to incorporate additional predictors and informed prior probabilities, contributing to enhanced prediction accuracy for rare diseases. It also allows predictions without these additional tests, providing additional information on whether a patient should undergo further biomarker testing before genetic testing.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Bias reduction and inference for electronic health record data under selection and phenotype misclassification: three case studies 93%
- A Stability-Enhanced Lasso Approach for Covariate Selection in Non-Linear Mixed Effect Model 93%
- A Double Machine Learning Approach for the Evaluation of COVID-19 Vaccine Effectiveness under the Test-Negative Design: Analysis of Québec Administrative Data 93%
Similar papers in this journal
- Bayesian Structural Time Series for Biomedical Sensor Data: A Flexible Modeling Framework for Evaluating Interventions 93%
- Model guided trait-specific co-expression network estimation as a new perspective for identifying molecular interactions and pathways 92%
- A time-series analysis of blood-based biomarkers within a 25-year longitudinal dolphin cohort. 92%
Similar papers in this journal
- Towards development of a statistical framework to evaluate myotonic dystrophy type 1 mRNA biomarkers in the context of a clinical trial 94%
- Future prevalence of type 2 diabetes – a comparative analysis of chronic disease projection methods 93%
- Time-to-event estimation of birth year prevalence trends: a method to enable investigating the etiology of childhood disorders including autism 93%
Similar papers in this journal
- Towards reduction in bias in epidemic curves due to outcome misclassification through Bayesian analysis of time-series of laboratory test results: Case study of COVID-19 in Alberta, Canada and Philadelphia, USA 93%
- External control arm analysis: an evaluation of propensity score approaches, G-computation, and doubly debiased machine learning 92%
- Prediction-powered Inference for Clinical Trials 92%
Similar papers in this journal
- TWO-SIGMA: a novel TWO-component SInGle cell Model-based Association method for single-cell RNA-seq data 94%
- Statistics to prioritize rare variants in family-based sequencing studies with disease subtypes 93%
- Case-Control Vs. Case-Only Estimates Of Gene-Environment Interactions With Common And Misclassified Clinical Diagosis 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.