Back

Leveraging Language Embeddings from EMA Surveys to Predict Perceived Social Isolation among Stroke Survivors

Liu, Y.; Wong, A.; Fong, M.; Metts, C.; Shi, Y.; Lee, S. I.

2025-07-17 health informatics
10.1101/2025.07.17.25331714 medRxiv
Show abstract

Perceived social isolation (PSI) significantly affects the emotional well-being of stroke survivors, necessitating effective monitoring and prediction for timely, targeted interventions. While Ecological Momentary Assessment (EMA) has been increasingly used to identify precursor characteristics of PSI, existing prediction methods rely on handcrafted features, which often fail to capture the semantic richness and contextual relationships among survey questions. In this study, we propose a novel approach to predict PSI by processing structured EMA data with language embeddings. A total of 11,802 EMA surveys were collected from 218 stroke survivors, the largest dataset of its kind in social isolation research for this population. Language embeddings were extracted from the structured EMA surveys using a pre-trained language model. These embeddings were then processed by training an autoencoder to generate compact latent representations, which were used for the downstream PSI prediction. Our findings show that the proposed approach achieves accurate PSI prediction, with a weighted F1 score of 0.84 and a weighted AUPRC of 0.92, outperforming traditional handcrafted features. Furthermore, by leveraging only three carefully selected questions, our method can optimize a trade-off between validity and usability. This study demonstrates an efficient method for real-time monitoring of psychosocial outcomes in stroke survivors, with potential implications for early intervention and personalized care.

Published in IEEE Journal of Biomedical and Health Informatics (predicted rank #1) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.