From Keywords to Context: Bridging Expert Insight and Language Models for Multidimensional Sleep Health Classification in Clinical Notes
Hussain, S.-A.; Calloway, A.; Sirrianni, J. W.; Fosler-Lussier, E.; Davenport, M.
Show abstract
Accurate detection of multidimensional sleep health (MSH) information from electronic health records (EHRs) is critical for improving clinical decision-making but remains challenging due to sparse documentation and class imbalance. This study investigates whether integrating expert-guided annotations and keyword-based heuristics with large language models (LLMs) enhances the extraction of nuanced MSH indicators from clinical narratives. Using a novel, expertly annotated dataset (NCH-Sleep), we trained and evaluated models to classify clinical notes across nine clinically relevant MSH categories. Our baseline model demonstrated substantial predictive capability using raw text alone. Incorporating manually annotated spans (oracle annotations) dramatically improved performance, highlighting the benefit of targeted expert guidance. Additionally, employing curated keyword annotations within varying context windows significantly enhanced model interpretability while retaining strong predictive accuracy. Through detailed bias analyses, we identified consistent performance across demographics and clinical settings, although specific disparities underscored the importance of balanced expert oversight. Our findings emphasize the value of expert-informed supervision and heuristic approaches in building scalable, interpretable clinical NLP systems for sleep health classification.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 95%
- Evaluating Anti-LGBTQIA+ Medical Bias in Large Language Models 92%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 92%
Similar papers in this journal
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 95%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 94%
- Using indication embeddings to represent patient health for drug safety studies 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.