Can NLP Detect Loneliness in Electronic Health Records? A Proof-of-Concept Study
Park, T.; Habibi, S.; Lowers, J.; Sarker, A.; Bozkurt, S.
Show abstract
Loneliness is clinically important but under-documented in electronic health records (EHRs), posing challenges for secondary use and computational phenotyping. This study evaluated whether natural language processing (NLP) methods can detect and classify loneliness severity from clinical notes. Patients with a loneliness survey (mild, moderate, severe) were identified, and notes within six months prior to the survey were retrieved. An expert-expanded lexicon was applied, and transformer models (RoBERTa, ClinicalBERT, Longformer) were fine-tuned for loneliness severity classification. Large language model-based summarization of social and psychiatric history was also tested as an alternative input representation. Performance was evaluated using accuracy, weighted-F1, and per-class F1. All models achieved modest accuracy (0.3 to 0.7), and struggled to identify severe loneliness, reflecting sparse and inconsistent documentation even among surveyed patients. While summarization marginally improved accuracy, gains primarily reflected mild predictions. Manual review of 100 social worker notes from severely lonely patients found explicit mentions of loneliness in only two cases, confirming that relevant documentation is exceedingly rare. These findings demonstrate that model performance is constrained by the sparse and inconsistent documentation of loneliness in EHRs, rather than by deficiencies in the modeling approach itself.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Use of unstructured text in prognostic clinical prediction models: a systematic review 93%
- Annotation-preserving machine translation of English corpora to validate Dutch clinical concept extraction tools 93%
- Beyond Metrics to Methods: A Scoping Review of Large Language Models for Detection of Social Drivers of Health in Clinical Notes 93%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Long COVID symptoms from Reddit: Characterizing post-COVID syndrome from patient reports 93%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 92%
- Evaluation of Patient-Level Retrieval from Electronic Health Record Data for a Cohort Discovery Task 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.