Internal and External Validation of Machine Learning Algorithms Versus FINDRISC for Incident Type 2 Diabetes: A Transparent, Explainable Benchmark Using SHAP
Amirian, P.; Zarpoosh, M.
Show abstract
BackgroundType 2 diabetes mellitus (T2DM) affects almost half a billion people, and the projected cost is $2.25 trillion by 2030; early detection strategies are needed for prevention. However, Finnish Diabetes Risk Score (FINDRISC) enables questionnaire-based risk assessment; its linearity and lack of interpretability limit predictive power and generalizability across diverse populations. Machine learning (ML) models are the next-generation prediction tools, but require rigorous benchmarking, external validation, and explainability to be clinically trusted. MethodsIn the current prospective cohort (n=9,171, 7.1 years follow-up), we compared six supervised ML models, three anomaly detectors, and a stacking ensemble against FINDRISC for T2DM incidence, using harmonized, calibrated pipelines and internal and external validation in US (NHANES) and PIMA Indian populations. External validations included reduced (7- and 3-variable) models, and explainability was assessed with SHAP. ResultsML models, particularly neural networks and stacking, achieved superior internal discrimination (ROC AUC up to 0.87 vs. FINDRISC 0.70), with stacking ensemble recall of 0.81. In reduced-variable external validations, ML models maintained robust performance (AUCs > 0.76), and strikingly, the isolation forest anomaly detector excelled in US data. Sensitivity analysis demonstrated that without laboratory data, FINDRISC still matches or exceeds ML, thereby preserving its practical role in non-laboratory settings. SHAP analysis consistently identified FBS, BMI, and age as main predictors, promoting interpretability. ConclusionsHarmonized ML models, when externally validated, substantially improve traditional risk scores for T2DM prediction, particularly when laboratory data are available. Transparent analytics and an open-source online calculator support global clinical deployment. This work substantially advances precision prevention in T2DM through explainable, portable prediction.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Unusual patterns of risk for diabetic retinopathy in the Mumbai slums: The Aditya Jyot Diabetic Retinopathy in Urban Mumbai Slums Study (AJ-DRUMSS) Report 3 91%
- Perceived risk of type 2 diabetes: Using linked genomic, clinical and questionnaire data to understand the potential use of genetic risk tools in British South Asians 90%
- Prediction models for post-discharge mortality among under-five children with suspected sepsis in Uganda: A multicohort analysis 90%
Similar papers in this journal
- Estimating Heritability of Glycaemic Response to Metformin using Nationwide Electronic Health Records and Population-Sized Pedigree 92%
- Subpopulation-specific Machine Learning Prognosis for Underrepresented Patients with Double Prioritized Bias Correction 92%
- Pretrained Patient Trajectories for Adverse Drug Event Prediction Using Common Data Model-based Electronic Health Records 91%
Similar papers in this journal
- Study Research Protocol for Phenome India-CSIR Health Cohort Knowledgebase (PI-CHeCK): A Prospective multi-modal follow-up study on a nationwide employee cohort. 91%
- Machine Classification of Methylomes in Cancer 90%
- A Novel Method for Handling Pre-Existing Conditions in Prediction Models for Covid-19 Death 89%
Similar papers in this journal
- Racial disparities in continuous glucose monitoring-based 60-min glucose predictions among people with type 1 diabetes 95%
- Machine learning-based equations for improved body composition estimation in Indian adults 92%
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 92%
Similar papers in this journal
- Assessing the transportability of clinical prediction models for cognitive impairment using causal models 91%
- Quantitative bias analysis in practice: Review of software for regression with unmeasured confounding 90%
- External control arm analysis: an evaluation of propensity score approaches, G-computation, and doubly debiased machine learning 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.