Large Language Model-Driven Evaluation of Medical Records Using MedCheckLLM
Schubert, M. C.; Wick, W.; Venkataramani, V.
Show abstract
Large Language Models (LLMs) offer potential in healthcare, especially in the evaluation of medical documents. This research introduces MedCheckLLM, a multi-step framework designed for the systematic assessment of medical records against established evidence-based guidelines, a process termed guideline-in-the-loop. By keeping the guidelines separate from the LLMs training data, this approach emphasizes validity, flexibility, and interpretability. Suggested evidence-based guidelines are externally accessed and fed back into the LLM for a evaluation. The method enables implementation of guideline updates and personalized protocols for specific patient groups without retraining. We applied MedCheckLLM to expert-validated simulated medical reports, focusing on headache diagnoses following International Headache Society guidelines. Findings revealed MedCheckLLM correctly extracted diagnoses, suggested appropriate guidelines, and accurately evaluated 87% of checklist items, with its evaluations aligning significantly with expert opinions. The system not only enhances healthcare quality assurance but also introduces a transparent and efficient means of applying LLMs in clinical settings. Future considerations must address privacy and ethical concerns in actual clinical scenarios.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Transformative potential of Large Language Models in data mining on Electronic Health Records. 95%
- Evaluating the impact on clinical task efficiency of a natural language processing algorithm for searching medical documents: Prospective crossover study 94%
- FHIR-DHP: A Standardized Clinical Data Harmonisation Pipeline for scalable AI application deployment 94%
Similar papers in this journal
- Temporal Relationship of Computed and Structured Diagnoses in Electronic Health Record Data 92%
- Automated abstraction of clinical parameters of multiple myeloma from real-world clinical notes using large language models 92%
- An Interpretable Risk Prediction Model for Healthcare with Pattern Attention 92%
Similar papers in this journal
- Dataset Documentation for Responsible AI: Analysis of Suitability and Usage for Health Datasets 95%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 94%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
Similar papers in this journal
- Transforming Estonian health data to the Observational Medical Outcomes Partnership (OMOP) Common Data Model: lessons learned 94%
- A Simple Electronic Medical Record System Designed for Research 93%
- Evaluation of Patient-Level Retrieval from Electronic Health Record Data for a Cohort Discovery Task 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.