Extraction and validation of patient housing and food insecurity status in a large electronic health records database using selective prediction and active learning
Swaminathan, A.; Kumar, W. M.; Lopez, I.; Tran, E.; Srivastava, U.; Wang, W.; Clary, M. K.; van Deursen, W. H.; Shaw, J. G.; Gevaert, O.; Chen, J. H.
Show abstract
ObjectiveInformation on patient social determinants of health is frequently recorded in unstructured clinical notes, making it inaccessible for researchers and policymakers. We aimed to extract and validate food and housing insecurity status on a large electronic health record-derived patient cohort by combining selective prediction and active learning. Materials and MethodsManually labeled charts selected via active learning were used to train L1-regularized logistic regression models to identify the presence of food insecurity (N=372, 42% event rate) and housing insecurity (N=559, 36% event rate) in clinical notes. In addition to validating predictions against labeled data, we further validated predictions on an additional unlabeled dataset through associative studies with demographic, clinical, and environmental variables with known associations with food and housing insecurity. ResultsThe food insecurity model had AUC=0.83, sensitivity=0.90, PPV=0.90, and undetermined rate=0.59 (n=149); the housing insecurity model had AUC=0.81, sensitivity=0.50, PPV=1, and undetermined rate=0.65 (n=224). Out of 4,337 unlabeled patients, the 395 (9%) patients predicted to have food insecurity were more likely to be Hispanic/Latino (48% vs 24%, p<0.001) and have diabetes (34% vs 12%), hypertension (43% vs 11%), and heart disease (12% vs 0.7%) (p<0.001 for all). DiscussionSelective prediction and active learning can facilitate efficient labeling of social determinants of health from unstructured EHR data to identify vulnerable populations and targets for healthcare system and policy intervention. ConclusionMachine learning can be used to extract high-fidelity information on patient food and housing insecurity status.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Real-Time Electronic Health Record Mortality Prediction During the COVID-19 Pandemic: A Prospective Cohort Study 95%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 95%
- Machine Learning Approaches for Electronic Health Records Phenotyping: A Methodical Review 94%
Similar papers in this journal
- A Deep Learning Method to Detect Opioid Prescription and Opioid Use Disorder from Electronic Health Records 93%
- Predicting Prognosis in COVID-19 Patients using Machine Learning and Readily Available Clinical Data 92%
- Predicting nutrition and environmental factors associated with female reproductive disorders using a knowledge graph and random forests 92%
Similar papers in this journal
- Characterization and Racial Stratification of Social Determinants of Health for Individuals with Type 2 Diabetes as Recorded in Electronic Health Records: Implications for Artificial Intelligence Development 93%
- Development and Application of Pharmacological Statin-Associated Muscle Symptoms Phenotyping Algorithms Using Structured and Unstructured Electronic Health Records Data 93%
- Modeling physician variability to prioritize relevant medical record information 92%
Similar papers in this journal
- A scoping review of fair machine learning techniques when using real-world data 94%
- Natural language processing for scalable feature engineering and ultra-high-dimensional confounding adjustment in healthcare database studies 94%
- A Deep Learning Approach for Transgender and Gender Diverse Patient Identification in Electronic Health Records 93%
Similar papers in this journal
- Score for Emergency Risk Prediction (SERP): An Interpretable Machine Learning AutoScore–Derived Triage Tool for Predicting Mortality after Emergency Admissions 94%
- Diagnostic Codes in AI prediction models and Label Leakage of Same-admission Clinical Outcomes 93%
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.