Approach to Machine Learning for Extraction of Real-World Data Variables from Electronic Health Records
Adamson, B. J.; Waskom, M.; Blarre, A.; Kelly, J.; Krismer, K.; Nemeth, S.; Gipetti, J.; Ritten, J.; Harrison, K.; Ho, G.; Linzmayer, R.; Bansal, T.; Wilkinson, S.; Amster, G.; Estola, E.; Benedum, C. M.; Fidyk, E.; Estevez, M.; Shapiro, W.; Cohen, A. B.
Show abstract
BackgroundAs artificial intelligence (AI) continues to advance with breakthroughs in natural language processing (NLP) and machine learning (ML), such as the development of models like OpenAIs ChatGPT, new opportunities are emerging for efficient curation of electronic health records (EHR) into real-world data (RWD) for evidence generation in oncology. Our objective is to describe the research and development of industry methods to promote transparency and explainability. MethodsWe applied NLP with ML techniques to train, validate, and test the extraction of information from unstructured documents (eg, clinician notes, radiology reports, lab reports, etc.) to output a set of structured variables required for RWD analysis. This research used a nationwide electronic health record (EHR)-derived database. Models were selected based on performance. Variables curated with an approach using ML extraction are those where the value is determined solely based on an ML model (ie, not confirmed by abstraction), which identifies key information from visit notes and documents. These models do not predict future events or infer missing information. ResultsWe developed an approach using NLP and ML for extraction of clinically meaningful information from unstructured EHR documents and found high performance of output variables compared with variables curated by manually abstracted data. These extraction methods resulted in research-ready variables including initial cancer diagnosis with date, advanced/metastatic diagnosis with date, disease stage, histology, smoking status, surgery status with date, biomarker test results with dates, and oral treatments with dates. ConclusionsNLP and ML enable the extraction of retrospective clinical data in EHR with speed and scalability to help researchers learn from the experience of every person with cancer.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Adoption of the OMOP CDM for Cancer Research using Real-world Data: Current Status and Opportunities 94%
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 94%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 93%
Similar papers in this journal
- Natural language inference for clinical registry curation 95%
- LCD Benchmark: Long Clinical Document Benchmark on Mortality Prediction for Language Models 95%
- Is One Run Enough? Reproducibility of Flagship Large Language Models Across Temperature and Reasoning Settings in Biomedical Text Processing 94%
Similar papers in this journal
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 94%
- EHR Foundation Models Improve Robustness in the Presence of Temporal Distribution Shift 94%
- Scalable Incident Detection via Natural Language Processing and Probabilistic Language Models 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.