Back

Can Large Language Models Reduce the Cost of Extracting Data from Electronic Health Records for Research?

Hagler, S.; Adibuzzaman, M.; McWeeney, S. K.; Cohen, A. M.

2026-01-11 health informatics
10.64898/2026.01.09.26343792 medRxiv
Show abstract

ObjectiveMuch medical data is only available in unstructured electronic health records (EHR). These data can be obtained through manual (human) extraction or programmatic natural language processing (NLP) methods. We estimate that NLP only becomes economically competitive with manual extraction when there are ~6500 EHRs records. We have found that there is interest from clinicians and researchers in using NLP on projects with fewer records. We examine whether a large language model (LLM) can be used to reduce the cost of NLP to make it economically competitive for such projects, and study the feasibility of such framework for accuracy. Materials and MethodsWe developed an NLP pipeline using an off-the-shelf open LLM to extract breast cancer ER, PR, and HER2 biomarker data. Pipeline development stopped when the prompts performances were competitive with manual extraction. The development time and extraction performance were compared to those for an existing rule-based (RB) NLP pipeline. The code for the extraction portion of the LLM pipeline is available at https://github.com/sehagler/llm_biomarker_extraction. ResultsThe LLM pipeline produced performance competitive with manual data extraction with a hands-on development time that was ~38% that of the RB pipeline. DiscussionLLMs exhibit lower hands-on development costs compared to standard NLP techniques, but require significant and potentially costly computation resources. ConclusionLLMs may potentially allow the economically competitive application of NLP to smaller projects if computation costs can be managed.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.