Back

A Transparent Four-Feature Logistic Model for Depression Screening in Assisted-Living Facilities

Mekulu, K.; Aqlan, F.; Yang, H.

2025-07-16 health informatics
10.1101/2025.07.14.25331539 medRxiv
Show abstract

Depression in older adults is both common and frequently underdiagnosed, especially in assisted-living communities, where it often co-occurs with mild cognitive impairment (MCI), creating a complex and vulnerable clinical landscape. Despite this urgency, scalable, interpretable, and easy-to-administer tools for early screening remain scarce. In this study, we introduce a transparent and lightweight AI-driven screening model that uses only four linguistic features extracted from brief conversational speech, to detect depression with high sensitivity. Trained on the DAIC-WOZ dataset and optimized for deployment in resource constrained settings, our model achieved strong discriminative performance (AUC = 0.760) with a clinically calibrated sensitivity of 92%. Beyond raw accuracy, the model offers insights into how affective language, syntactic complexity, and latent semantic content relate to psychological states. Notably, one semantic feature derived from transformer embeddings, emb_1, appears to capture deeper emotional or cognitive tension not directly expressed through lexical negativity. We propose this component as a potential digital biomarker of cognitive-affective strain, warranting further longitudinal study. Our approach outperforms many more complex models in the literature, yet remains simple enough for real-time, on-device use, marking a step forward in making mental health AI both interpretable and clinically actionable.

Published in Frontiers in Digital Health (predicted rank #2) · training set

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.