Can Large Language Models Reduce the Cost of Extracting Data from Electronic Health Records for Research?
Hagler, S.; Adibuzzaman, M.; McWeeney, S. K.; Cohen, A. M.
Show abstract
ObjectiveMuch medical data is only available in unstructured electronic health records (EHR). These data can be obtained through manual (human) extraction or programmatic natural language processing (NLP) methods. We estimate that NLP only becomes economically competitive with manual extraction when there are ~6500 EHRs records. We have found that there is interest from clinicians and researchers in using NLP on projects with fewer records. We examine whether a large language model (LLM) can be used to reduce the cost of NLP to make it economically competitive for such projects, and study the feasibility of such framework for accuracy. Materials and MethodsWe developed an NLP pipeline using an off-the-shelf open LLM to extract breast cancer ER, PR, and HER2 biomarker data. Pipeline development stopped when the prompts performances were competitive with manual extraction. The development time and extraction performance were compared to those for an existing rule-based (RB) NLP pipeline. The code for the extraction portion of the LLM pipeline is available at https://github.com/sehagler/llm_biomarker_extraction. ResultsThe LLM pipeline produced performance competitive with manual data extraction with a hands-on development time that was ~38% that of the RB pipeline. DiscussionLLMs exhibit lower hands-on development costs compared to standard NLP techniques, but require significant and potentially costly computation resources. ConclusionLLMs may potentially allow the economically competitive application of NLP to smaller projects if computation costs can be managed.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 94%
- Ontology-based expansion of virtual gene panels to improve diagnostic efficiency for rare genetic diseases 93%
- Towards a Clinically-based Common Coordinate Framework for the Human Gut Cell Atlas - The Gut Models 93%
Similar papers in this journal
Similar papers in this journal
- A deep learning workflow for quantification of Micronuclei in DNA damage studies in cultured cancer cell lines: a proof of principle investigation 93%
- Predicting gene and protein expression levels from DNA and protein sequences with Perceiver 93%
- A standardized analytics pipeline for reliable and rapid development and validation of prediction models using observational health data 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.