Identifying Prediabetes in Canadian Populations Using Machine Learning
Lu, K.; Sheth, P.; Zhou, Z. L.; Kazari, K.; Guergachi, A.; Keshavjee, K.; Noaeen, M.; Shakeri, Z.
Show abstract
Prediabetes is a critical health condition characterized by elevated blood glucose levels that fall below the threshold for Type 2 diabetes (T2D) diagnosis. Accurate identification of prediabetes is essential to forestall the progression to T2D among at-risk individuals. This study aims to pinpoint the most effective machine learning (ML) model for prediabetes prediction and to elucidate the key biological variables critical for distinguishing individuals with prediabetes. Utilizing data from the Canadian Primary Care Sentinel Surveillance Network (CPCSSN), our analysis included 6,414 participants identified as either nondiabetic or prediabetic. A rigorous selection process led to the identification of ten variables for the study, informed by literature review, data completeness, and the evaluation of collinearity. Our comparative analysis of seven ML models revealed that the Deep Neural Network (DNN), enhanced with early stop regularization, outshined others by achieving a recall rate of 60%. This models performance underscores its potential in effectively identifying prediabetic individuals, showcasing the strategic integration of ML in healthcare. While the model reflects a significant advancement in prediabetes prediction, it also opens avenues for further research to refine prediction accuracy, possibly by integrating novel biological markers or exploring alternative modeling techniques. The results of our work represent a pivotal step forward in the early detection of prediabetes, contributing significantly to preventive healthcare measures and the broader fight against the global epidemic of Type 2 diabetes.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Clinical interpretation of machine learning models for prediction of diabetic complications using electronic health records 97%
- Trajectories: a framework for detecting temporal clinical event sequences from health data standardized to the OMOP Common Data Model 92%
- Characterizing subgroup performance of probabilistic phenotype algorithms within older adults: A case study for dementia, mild cognitive impairment, and Alzheimer’s and Parkinson’s diseases 92%
Similar papers in this journal
- Racial disparities in continuous glucose monitoring-based 60-min glucose predictions among people with type 1 diabetes 97%
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 92%
- Automated Image Transcription for Perinatal Blood Pressure Monitoring Using Mobile Health Technology 91%
Similar papers in this journal
- Predicting long-term Type 2 Diabetes with Support Vector Machine using Oral Glucose Tolerance Test 96%
- Determinants of glycemic control among persons living with type 2 diabetes mellitus attending a district hospital in Ghana 94%
- Predictors of glycemic control, quality of life and diabetes self-management of patients with diabetes mellitus at a tertiary hospital in Ghana 94%
Similar papers in this journal
- Testing the phenotypic decanalization hypothesis: social determinants of hyperglycemia and type 2 diabetes in adult urban Argentinian population 93%
- Leveraging Large Language Models to Analyze Continuous Glucose Monitoring Data: A Case Study 93%
- Machine learning for classifying chronic kidney disease and predicting creatinine levels using at-home measurements 93%
Similar papers in this journal
- Prediction of Sepsis Mortality in ICU Patients Using Machine Learning Methods 92%
- Optimized Feature Selection and Advanced Machine Learning for Stroke Risk Prediction in Revascularized Coronary Artery Disease Patients 91%
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.