A Scalable Method for Validated Data Extraction from Electronic Health Records with Large Language Models
Stuhlmiller, T. J.; Rabe, A.; Rapp, J.; Manasco, P.; Awawda, A.; Kouser, H.; Salamon, H.; Chuyka, D.; Mahoney, W.; Wong, K. K.; Kramer, G. A.; Shapiro, M. A.
Show abstract
PurposeExtracting and structuring relevant clinical information from electronic health records (EHRs) remains a challenge due to the heterogeneity of systems, documents, and documentation practices. Large Language Models (LLMs) provide an approach to processing semi-structured and unstructured EHR data, enabling classification, extraction, and standardization. MethodsMedical documents are processed through a structured data pipeline to generate normalized FHIR data. Unstructured data undergoes preprocessing, including optical character recognition, document parsing, text chunking and embedding. Embedding enables search and classification which facilitate document retrieval for extraction. LLMs perform named entity recognition and relation extraction, with outputs mapped to FHIR R4 and OMOP and harmonized with pre-structured data for interoperability. Model performance is evaluated through human validation and automated consistency checks. Iterative refinement, error analysis, and standardized schema selection optimize use for analytics and downstream workflows. ResultsAn LLM-based schema for medication extraction, validated on 34 patients, achieved [~]95% accuracy and F1-score across 7 data fields for 5,789 extracted medications. Deployment across 11,115 patients extracted 2.6 million medication records, increasing total medications by 27% and distinct drug ingredients by 31% over structured data in the EHR. LLM extraction increased oncology medications by 60%, distinct oncology therapies by 64%, and the number of patients with structured oncology medication data by 33%. The LLM enhanced data completeness, improving availability of indication for prescription (61% vs. 31%) and discontinuation reason (17% vs. 0%), outperforming pre-structured data in key clinical variables. ConclusionAn LLM-powered extraction process that employs embeddings, machine learning classification, schema-based extraction, and mapping of extracted information to healthcare data standards, achieves a significant gain in clinically relevant information over pre-structured data available in the EHR.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Design and Implementation of an End-to-End AI-Driven Colonoscopy Recall Workflow at Scale 96%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 95%
- Transforming Estonian health data to the Observational Medical Outcomes Partnership (OMOP) Common Data Model: lessons learned 95%
Similar papers in this journal
- EHR-QC: A streamlined pipeline for automated electronic health records standardisation and preprocessing to predict clinical outcomes 96%
- Biomedical Text Normalization through Generative Modeling 95%
- Development of a Post-Acute Sequelae of COVID-19 (PASC) Symptom Lexicon Using Electronic Health Record Clinical Notes 95%
Similar papers in this journal
- Evaluating the impact on clinical task efficiency of a natural language processing algorithm for searching medical documents: Prospective crossover study 94%
- Transformative potential of Large Language Models in data mining on Electronic Health Records. 94%
- FHIR-DHP: A Standardized Clinical Data Harmonisation Pipeline for scalable AI application deployment 94%
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 96%
- Adoption of the OMOP CDM for Cancer Research using Real-world Data: Current Status and Opportunities 93%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.