MedSDoH: A Rule-Based System for Extracting Social Determinants of Health from Multi-site EHRs Based on the OHNLP Framework
Ahn, J.; Fu, S.; Palacios, D. M.; Jeong, H.-H.; Wang, L.; Swartz, M. C.; Tosur, M.; Redondo, M. J.; Wu, X.; Yue, Z.; Kakadiaris, A.; Wang, N.; Li, Z.; Huang, M.; Wen, A.; Harris, D.; Wang, Y.; Kwak, M. J.; Liu, Z.; Liu, H.
Show abstract
ObjectiveSocial Determinants of Health (SDoH) are critical to patient care and population health. Despite their importance, SDoH information is frequently embedded within unstructured clinical text such as patient-reported information or social worker notes, which limits its use on clinical decision-making and resource allocation. Although transformer-based models represent the current state of the art, their scalability, computational requirements, and limited transparency pose barriers to large-scale multi-site clinical implementation. In this context, rule-based NLP systems remain valuable, particularly when explainability, reproducibility, and rapid customization are essential. MethodsMedSDoH was developed within the Open Health Natural Language Processing (OHNLP) Framework using literature-derived SDoH resources, standardized domain definitions, and expert-curated rulesets. Large language models (LLMs) were used during development to assist with rule generation and lexicon expansion. Rules were iteratively refined against a gold-standard annotated corpus from two health systems and then evaluated on independent datasets. ResultThe final system included 942 regular expression rules spanning 22 SDoH domains. On validation on two external datasets, MedSDoH demonstrated generalizability and comparable performance across sites. The system has been made publicly available so research community can collaboratively contribute to the maintenance and extension through disease- or site-specific adaptations. ConclusionMedSDoH is a computationally efficient and open-source system for large-scale SDoH extraction from clinical text. It is well-suited for multi-site adaptation and deployment in resource-constrained settings.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 95%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 95%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 95%
Similar papers in this journal
- EHR-QC: A streamlined pipeline for automated electronic health records standardisation and preprocessing to predict clinical outcomes 96%
- Development of a Post-Acute Sequelae of COVID-19 (PASC) Symptom Lexicon Using Electronic Health Record Clinical Notes 96%
- ConceptWAS: a high-throughput method for early identification of COVID-19 presenting symptoms 95%
Similar papers in this journal
- Transforming Estonian health data to the Observational Medical Outcomes Partnership (OMOP) Common Data Model: lessons learned 95%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 95%
- Development of a COVID-19 Application Ontology for the ACT Network 95%
Similar papers in this journal
- Evaluating Anti-LGBTQIA+ Medical Bias in Large Language Models 93%
- From months to minutes: creating Hyperion, a novel data management system expediting data insights for oncology research and patient care 93%
- Raising awareness of potential biases in medical machine learning: Experience from a Datathon 93%
Similar papers in this journal
- Evaluating the impact on clinical task efficiency of a natural language processing algorithm for searching medical documents: Prospective crossover study 95%
- Extracting social determinants of health from electronic health records: development and comparison of rule-based and large language models-based methods 95%
- Transformative potential of Large Language Models in data mining on Electronic Health Records. 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.