Developing an open-source framework for LLM evaluation of patients using EHR clinical documentation; performance of LLMs relative to medical professionals
Barrett, L.; Joshi, N.; North, A. S.; Dimitrov, L.; Maughan, E. F.; Ross, T.; Pankhania, R.; Paramjothy, K.; Minty, I.; Farache-Trajano, L.; Smith, S. L.; Mason, K. A.; Bhargava, E. K.; Donnelly, C.; Fatoum, H.; Padiyar, A.; Kader, Z.; Chan, C. H. K.; Schilder, A. G.; Mehta, N.
Show abstract
Background: Large language models (LLMs) have shown increasing capability in medical knowledge tasks, yet how they perform in extracting structured clinical information from real-world clinical documentation remains uncertain. We evaluated the performance of LLMs relative to medical professionals in extracting SNOMED-coded clinical information from openly available Ear, Nose and Throat (ENT) EHRs from MTSamples, examining both reliability and accuracy metrics. Methods: We evaluated the performance of seven LLMs (including GPT-4o, Claude 3.5, Gemini 1.5 Pro, Gemma 3 and three LLAMA variants) against annotations from fourteen medical professionals who served as both study authors and data annotators. Each annotator independently extracted seven categories of clinical information from 98 publicly available ENT clinical documents: socio-demographics, symptoms, signs, diagnoses, treatments, risk factors, and test results. Standardised medical terminology was enforced through SNOMED-CT code assignment, enabling standardised comparison through Cohen's Kappa. We employed Bayesian hierarchical modelling to test non-inferiority of medic-LLM agreement compared to medic-medic agreement, using Beta distributed likelihood functions with weakly informative priors. Non-inferiority margins of 0.05, 0.10, and 0.15 were assessed with 95% posterior probability thresholds. Results: Cohen's Kappa for inter-rater reliability was 0.752 (95% CI: 0.710 - 0.794) between medical professionals and 0.391 (95% CI: 0.362-0.420) between LLMs and medical professionals. Bayesian analysis showed medic-medic agreement (posterior mean 0.813, 95% CI: 0.755-0.860) exceeded medic-LLM agreement (0.659, 95% CI: 0.633-0.684) by 0.154 (95% CI: 0.091-0.209). Non-inferiority was rejected at all tested margins (delta = 0.05, 0.10, 0.15). Agreement varied by clinical category, with smallest differences for test results and largest for diagnoses. GPT-4o achieved 97.0% precision and 84.9% recall, with a 7.5% false positive rate. Conclusions: Current LLMs do not achieve inter-rater reliability levels comparable to medical professionals in clinical information extraction from ENT documentation. These findings provide evidence-based guidance for LLM deployment in clinical documentation workflows, suggesting they are best suited for initial extraction with human verification rather than autonomous operation.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 94%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 93%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 92%
Similar papers in this journal
- Generative Large Language Models in Electronic Health Records for Patient Care Since 2023: A Systematic Review 95%
- What Do Clinicians Edit in Ambient AI-Drafted Clinical Documentation? A Qualitative Content Analysis 94%
- Use of unstructured text in prognostic clinical prediction models: a systematic review 93%
Similar papers in this journal
- Comparison of local large language models for extraction of signs and symptoms data from electronic health records 96%
- Clinical code sets and the problem of redundancy in code set repositories 93%
- tbiExtractor: A framework for Extracting Traumatic Brain Injury Common Data Elements from Radiology Reports 93%
Similar papers in this journal
- Extraction of Crohn's Disease Clinical Phenotypes from Clinical Text Using Natural Language Processing 94%
- Systematic Review of Large Language Models for Patient Care: Current Applications and Challenges 93%
- Achieving Inclusive Healthcare through Integrating Education and Research with AI and Personalized Curricula 89%
Similar papers in this journal
- Evaluating the impact on clinical task efficiency of a natural language processing algorithm for searching medical documents: Prospective crossover study 95%
- Transformative potential of Large Language Models in data mining on Electronic Health Records. 94%
- Extracting social determinants of health from electronic health records: development and comparison of rule-based and large language models-based methods 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.