Predicting Diabetes in Canadian Adults Using Machine Learning
Esser, K.; Duong, M.; Kain, K.; Tran, S.; Sadeghi, A.; Guergachi, A.; Keshavjee, K.; Noaeen, M.; Shakeri, Z.
Show abstract
Rising diabetes rates have led to increased health-care costs and health complications. An estimated half of diabetes cases remain undiagnosed. Early and accurate diagnosis is crucial to mitigate disease progression and associated risks. This study addresses the challenge of predicting diabetes prevalence in Canadian adults by employing machine learning (ML) techniques to primary care data. We leveraged the Canadian Primary Care Sentinel Surveillance Network (CPCSSN), Canadas premier multi-disease electronic medical record surveillance system, and developed and tuned seven ML classification models to predict the likelihood of diabetes. The models were tested and validated, focusing on clinical patient characteristics influential in predicting diabetes. We found XGBoost performed best out of all the models, with an AUC of 92%. The most important features contributing to model prediction were HbA1c, LDL, and hypertension medication. Our research aims to aid healthcare professionals in early diagnosis and to identify key characteristics for targeted interventions. This study contributes to an understanding of how ML can enhance public health planning and reduce healthcare system burdens.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Racial disparities in continuous glucose monitoring-based 60-min glucose predictions among people with type 1 diabetes 97%
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 93%
- Automated Image Transcription for Perinatal Blood Pressure Monitoring Using Mobile Health Technology 93%
Similar papers in this journal
- Clinical interpretation of machine learning models for prediction of diabetic complications using electronic health records 97%
- Trajectories: a framework for detecting temporal clinical event sequences from health data standardized to the OMOP Common Data Model 94%
- Characterizing subgroup performance of probabilistic phenotype algorithms within older adults: A case study for dementia, mild cognitive impairment, and Alzheimer’s and Parkinson’s diseases 93%
Similar papers in this journal
- Predicting long-term Type 2 Diabetes with Support Vector Machine using Oral Glucose Tolerance Test 96%
- The Impact of Clinical Audits on Improving the Effectiveness of Type 2 Diabetes Mellitus (T2DM) CARE in Primary Health Centers. A Comprehensive Pre-post analysis through Multi-layered Intervention: The ICAE-DM CARE study protocol 95%
- Demographic and socioeconomic determinants of access to care: A subgroup disparity analysis using new equity-focused measurements 94%
Similar papers in this journal
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 91%
- Using Artificial Intelligence to Learn Optimal Regimen Plan for Alzheimer’s Disease 91%
- Causal modeling of chronic kidney disease in a participatory framework for informing the inclusion of social drivers in health algorithms 91%
Similar papers in this journal
- Prediction of Sepsis Mortality in ICU Patients Using Machine Learning Methods 93%
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 93%
- Optimized Feature Selection and Advanced Machine Learning for Stroke Risk Prediction in Revascularized Coronary Artery Disease Patients 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.