Leveraging Large Language Models for Digital Phenotyping: Detecting Depressive State Changes for Patients with Depressive Episodes
Yuan, Y.; Gao, Y.; Moen, H.; isometsä, e.; Marttinen, P.; Aledavood, T.
Show abstract
Digital phenotyping, which takes advantage of data continuously gathered from smartphones and wearable devices, offers promising avenues for real-time monitoring and mental health analysis. This approach holds promise for improving early detection and personalized care in mood disorders by enabling clinicians to proactively respond to significant changes before symptoms worsen. However, the complexity and heterogeneity of digital phenotyping data pose significant modeling challenges. Recent advances in large language models (LLMs) suggest their potential to generalize across diverse tasks with minimal labeled data, making them a promising alternative for analyzing data from digital phenotyping studies. However, the extent of usability of these methods for digital phenotyping studies is not yet well understood. In this study, we evaluate the potential of LLMs in analyzing digital phenotyping data to predict changes in depression severity among individuals experiencing major depressive episodes. We evaluate several in-context learning and fine-tuning strategies, and find that both few-shot prompted LLMs and fine-tuned models outperform traditional machine learning baselines trained on the same set of input features. Moreover, we compare two fine-tuning approaches (fine-tuning only the embedding layer versus parameter-efficient fine-tuning using QLoRA) and find that fine-tuning only the embedding layer significantly improves performance compared to QLoRA fine-tuning. These results highlight the capability of LLMs to process and integrate heterogeneous behavioral data, promising their application for digital phenotyping and mental health research. While our findings highlight the potential in using LLMs in mental health monitoring, their black-box nature and risk of replicating data biases highlight the need for clinical oversight and validation in real-world practice.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Labeling self-tracked menstrual health records with hidden semi-Markov models 93%
- Evaluating Explanations from AI Algorithms for Clinical Decision-Making: A Social Science-based Approach 93%
- A Transformer-Based Model Trained on Large Scale Claims Data for Prediction of Severe COVID-19 Disease Progression 93%
Similar papers in this journal
- Evaluating and mitigating unfairness in multimodal remote mental health assessments 94%
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 94%
- Emulation of epidemics via Bluetooth-based virtual safe virus spread: experimental setup, software, and data 94%
Similar papers in this journal
- High-Resolution Digital Phenotypes from Consumer Wearables Enhance Prediction of Cardiometabolic Risk Markers 92%
- One LLM is not Enough: Harnessing the Power of Ensemble Learning for Medical Question Answering 92%
- Quantified Flu: an individual-centered approach to gaining sickness-related insights from wearable data 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.