Back

Internal and External Validation of Machine Learning Algorithms Versus FINDRISC for Incident Type 2 Diabetes: A Transparent, Explainable Benchmark Using SHAP

Amirian, P.; Zarpoosh, M.

2025-09-07 endocrinology
10.1101/2025.09.05.25335151 medRxiv
Show abstract

BackgroundType 2 diabetes mellitus (T2DM) affects almost half a billion people, and the projected cost is $2.25 trillion by 2030; early detection strategies are needed for prevention. However, Finnish Diabetes Risk Score (FINDRISC) enables questionnaire-based risk assessment; its linearity and lack of interpretability limit predictive power and generalizability across diverse populations. Machine learning (ML) models are the next-generation prediction tools, but require rigorous benchmarking, external validation, and explainability to be clinically trusted. MethodsIn the current prospective cohort (n=9,171, 7.1 years follow-up), we compared six supervised ML models, three anomaly detectors, and a stacking ensemble against FINDRISC for T2DM incidence, using harmonized, calibrated pipelines and internal and external validation in US (NHANES) and PIMA Indian populations. External validations included reduced (7- and 3-variable) models, and explainability was assessed with SHAP. ResultsML models, particularly neural networks and stacking, achieved superior internal discrimination (ROC AUC up to 0.87 vs. FINDRISC 0.70), with stacking ensemble recall of 0.81. In reduced-variable external validations, ML models maintained robust performance (AUCs > 0.76), and strikingly, the isolation forest anomaly detector excelled in US data. Sensitivity analysis demonstrated that without laboratory data, FINDRISC still matches or exceeds ML, thereby preserving its practical role in non-laboratory settings. SHAP analysis consistently identified FBS, BMI, and age as main predictors, promoting interpretability. ConclusionsHarmonized ML models, when externally validated, substantially improve traditional risk scores for T2DM prediction, particularly when laboratory data are available. Transparent analytics and an open-source online calculator support global clinical deployment. This work substantially advances precision prevention in T2DM through explainable, portable prediction.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.