Wearable Prompt: In-Context Learning for Depression and Anxiety Prediction from Consumer Smart Ring Metrics
Azadifar, S.; Sameh, A.; Niemela, M.; Farrahi, V.
Show abstract
Large language models provide a promising framework for wearable-based health prediction by converting structured physiological and behavioral measurements into natural-language prompts. In this paper, we investigate whether pre-trained lightweight open-weight LLMs can predict depression and anxiety symptoms from short-horizon consumer wearable data. Using 4-8 days of Oura Ring data from 1,285 participants in the Northern Finland Birth Cohort 1986, we convert activity, sleep, heart rate, heart rate variability, demographic, and anthropometric measurements into structured prompts. We evaluate Llama 3.1, BioMistral, and Qwen 2.5 under zero-shot, rule-based, and few-shot in-context learning settings. To contextualize LLM performance, we compare them against machine learning models and recurrent neural networks. Our results show that prompt design is critical for LLM-based wearable inference. Zero-shot LLMs achieve high accuracy but largely predict the majority class, failing to identify participants with depression and anxiety symptoms. In contrast, few-shot prompting substantially improves positiveclass detection. Llama 3.1 with four in-context examples achieves the strongest performance, with 0.92 accuracy, 0.82 macro-F1, and 0.69 F1 for the positive class, among evaluated models. These findings suggest that lightweight LLMs can use in-context examples to better interpret structured wearable summaries and possibly provide a scalable direction for mental health prediction from consumer wearable data in combination with pre-trained LLMs.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Labeling self-tracked menstrual health records with hidden semi-Markov models 91%
- A Transformer-Based Model Trained on Large Scale Claims Data for Prediction of Severe COVID-19 Disease Progression 90%
- Leveraging Language Embeddings from EMA Surveys to Predict Perceived Social Isolation among Stroke Survivors 90%
Similar papers in this journal
- Listening to mental health crisis needs at scale: using Natural Language Processing to understand and evaluate a mental health crisis text messaging service 90%
- Using wearable and nearable devices in telerehabilitation for COPD: A review of digital endpoints in home-based programs 89%
- Remote digital measurement of visual and auditory markers of Major Depressive Disorder severity and treatment response. 87%
Similar papers in this journal
- Population Analysis Of Mortality Risk: Predictive Models Using Motion Sensors For 100,000 Participants In The UK Biobank National Cohort 92%
- Evaluating and mitigating unfairness in multimodal remote mental health assessments 91%
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 90%
Similar papers in this journal
- High-Resolution Digital Phenotypes from Consumer Wearables Enhance Prediction of Cardiometabolic Risk Markers 95%
- Quantified Flu: an individual-centered approach to gaining sickness-related insights from wearable data 91%
- Accurately Differentiating COVID-19, Other Viral Infection, and Healthy Individuals Using Multimodal Features via Late Fusion Learning 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.