Large language models for accurate disease detection in electronic health records
Burgisser, N.; Chalot, E.; Mehouachi, S.; Buclin, C. P.; Lauper, K.; Courvoisier, D. S.; Mongin, D.
Show abstract
ImportanceThe use of large language models (LLMs) in medicine is increasing, with potential applications in electronic health records (EHR) to create patient cohorts or identify patients who meet clinical trial recruitment criteria. However, significant barriers remain, including the extensive computer resources required, lack of performance evaluation, and challenges in implementation. ObjectiveThis study aims to propose and test a framework to detect disease diagnosis using a recent light LLM on French-language EHR documents. Specifically, it focuses on detecting gout ("goutte" in French), a ubiquitous French term that have multiple meanings beyond the disease. The study will compare the performance of the LLM-based framework with traditional natural language processing techniques and test its dependence on the parameter used. DesignThe framework was developed using a training and testing set of 700 paragraphs assessing "gout", issued from a random selection of retrospective EHR documents. All paragraphs were manually reviewed and classified by two health-care professionals (HCP) into disease (true gout) and non-disease (gold standard). The LLMs accuracy was tested using few-shot and chain-of-thought prompting and compared to a regular expression (regex)-based method, focusing on the effects of model parameters and prompt structure. The framework was further validated on 600 paragraphs assessing "Calcium Pyrophosphate Deposition Disease (CPPD)". SettingThe documents were sampled from the electronic health-records of a tertiary university hospital in Geneva, Switzerland. ParticipantsAdults over 18 years of age. ExposureMetas Llama 3 8B LLM or traditional method, against a gold standard. Main Outcomes and MeasuresPositive and negative predictive value, as well as accuracy of tested models. ResultsThe LLM-based algorithm outperformed the regex method, achieving a 92.7% [88.7-95.4%] positive predictive value, a 96.6% [94.6-97.8%] negative predictive value, and an accuracy of 95.4% [93.6-96.7%] for gout. In the validation set on CPPD, accuracy was 94.1% [90.2-97.6%]. The LLM framework performed well over a wide range of parameter values. Conclusions and RelevanceLLMs were able to accurately detect disease diagnoses from EHRs, even in non-English languages. They could facilitate creating large disease registries in any language, improving disease care assessment and patient recruitment for clinical trials. Key pointsO_ST_ABSQuestionC_ST_ABSHow accurate and efficient are large language models (LLMs) in detecting diseases from unstructured electronic health records (EHR) text compared to traditional natural language processing techniques? FindingsThis study proposes a framework based on Metas Llama 3 8B, a recent public LLM, outperforming traditional natural language processing techniques in detecting gout and calcium pyrophosphate deposition disease in unstructured text. It achieves high positive and negative predictive values and accuracy. Performance was robust over a wide range of parameters. MeaningThe proposed framework can ease the use of LLMs in effectively detecting disease in EHR data for various clinical applications.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Comparison of local large language models for extraction of signs and symptoms data from electronic health records 94%
- A method for rapid machine learning development for data mining with Doctor-In-The-Loop 93%
- Towards a COVID-19 symptom triad: The importance of symptom constellations in the SARS-CoV-2 pandemic 93%
Similar papers in this journal
- Transformative potential of Large Language Models in data mining on Electronic Health Records. 95%
- Is the quality of hospital EHR data sufficient to evidence its ICHOM outcomes performance in heart failure? A pilot evaluation 94%
- Evaluating the impact on clinical task efficiency of a natural language processing algorithm for searching medical documents: Prospective crossover study 94%
Similar papers in this journal
- Streamlining Intersectoral Provision of Real-World Health Data: A Service Platform for Improved Clinical Research and Patient Care 92%
- Machine Learning-based Clinical Decision Support for Infection Risk Prediction 91%
- Development and Validation of an Interpretable 3-day Intensive Care Unit Readmission Prediction Model Using Explainable Boosting Machines 91%
Similar papers in this journal
- Evaluation of a clinical decision support system for detection of patients at risk after kidney transplantation 94%
- Machine Learning Based Clinical Decision Support System for Early COVID-19 Mortality Prediction 91%
- AI-Driven Early Detection of Severe Influenza in Jiangsu, China: A Deep Learning Model Validated Through The Design of Multi-Center Clinical Trials and Prospective Real-World Deployment 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.