Back

Predicting Depression Among Canadians At-Risk or Living with Diabetes Using Machine Learning

Samsel, K.; Tiwana, A.; Ali, S.; Sadeghi, A.; Guergachi, A.; Keshavjee, K.; Noaeen, M.; Shakeri, Z.

2024-02-07 health informatics
10.1101/2024.02.03.24302303 medRxiv
Show abstract

Depression often goes unrecognized in individuals at risk or living with diabetes, presenting considerable challenges for primary care clinicians. Although large language models and other foundation model approaches are drawing significant attention, we systematically compared six established machine learning algorithms-Logistic Regression, Random Forest, AdaBoost, XGBoost, Naive Bayes, and Artificial Neural Networks-chosen for their reliability, interpretability, and feasibility in everyday clinical settings. By benchmarking their performance under real-world constraints, we identified key factors linked to depression risk in diabetes care, including patient sex, age, osteoarthritis, hemoglobin A1c, and body mass index. Although incomplete demographic information and potential label bias limited predictive power, our results demonstrate that a diverse set of clinical features might help pinpoint high-risk patients. They also indicate a need for longitudinal follow-up and richer clinical data to enhance model accuracy. As a practical benchmark for both clinicians and data scientists, this work suggests that machine learning-based risk stratification can improve early detection of depression and inform targeted interventions in diabetic populations.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.