Back

Quantifying Academic Risk Factors for Student Depression Using WHO Frameworks: Odds Ratios, SHAP Explainability, and Tipping Point Analysis

Ahmed, T.; Asif, M. R. A.

2026-08-11 psychiatry and clinical psychology
10.64898/2026.08.09.26360019 medRxiv
Show abstract

Depression has become a serious concern for students worldwide. Aligned with the WHO Helping Adolescents Thrive (HAT) Guidelines and the Social Determinants of Health (SDoH) model, this study isolates five academically relevant factors academic pressure, work/study hours, study satisfaction, sleep duration, and financial stress from a dataset of 27,880 university students in India and quantifies their associations with depression. Unlike prior work that maximises classification accuracy, this study prioritises interpretability: logistic regression provides odds ratios (OR) with 95% confidence intervals, Random Forest (RF) and XGBoost rank predictors by feature importance, and SHAP (SHapley Additive exPlanations) values extend the analysis to individual-level risk explanation. SMOTE oversampling was applied exclusively to the training set, and performance was evaluated on the original imbalanced test set (n = 5,576). Both ensemble models achieve approximately 77-78% accuracy and an AUC of 0.845, confirmed by 5-fold pipeline cross-validation (CV AUC [~] 0.843). Academic pressure is the dominant risk factor (OR = 2.271; RF importance = 0.481; mean |SHAP| = 0.174), while study satisfaction (OR = 0.796) and sleep duration (OR = 0.835) are protective. The RF model yields a tipping point at academic pressure > 4.02, and interaction plots reveal how depression risk is amplified by low sleep, high financial stress, and extended study hours. These findings provide data-driven thresholds aligned with WHO-endorsed modifiable determinants to support early detection and institutional counselling.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.