Predictive Modeling and Deep Phenotyping of Obstructive Sleep Apnea and Associated Comorbidities through Natural Language Processing and Large Language Models
Ahmed, A.; Rispoli, A.; Wasieloski, C.; Khurram, I.; Zamora-Resendiz, R.; Morrow, D.; Dong, A.; Crivelli, S.
Show abstract
Obstructive Sleep Apnea (OSA) is a prevalent sleep disorder associated with serious health conditions. This project utilized large language models (LLMs) to develop lexicons for OSA sub-phenotypes. Our study found that LLMs can identify informative lexicons for OSA sub-phenotyping in simple patient cohorts, achieving wAUC scores of 0.9 or slightly higher. Among the six models studied, BioClinical BERT and BlueBERT outperformed the rest. Additionally, the developed lexicons exhibited some utility in predicting mortality risk (wAUC score of 0.86) and hospital readmission (wAUC score of 0.72). This work demonstrates the potential benefits of incorporating LLMs into healthcare. Data and Code AvailabilityThis paper uses the MIMIC-IV dataset (Johnson et al., 2023a), which is available on the PhysioNet repository (Johnson et al., 2023b). We plan to make the source code publicly available in the future. Institutional Review Board (IRB)This research does not require IRB approval.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 93%
- Enriching Representation Learning Using 53 Million Patient Notes through Human Phenotype Ontology Embedding 92%
- Uncertainty in Deep Learning for EEG under Dataset Shifts 91%
Similar papers in this journal
Similar papers in this journal
- Evaluating Knowledge Fusion Models on Detecting Adverse Drug Events in Text 92%
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 91%
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.