Evaluating ChatGPT in Information Extraction: A Case Study of Extracting Cognitive Exam Dates and Scores
Jethani, N.; Jones, S.; Genes, N.; Major, V. J.; Jaffe, I. S.; Cardillo, A. B.; Heilenbach, N.; Ali, N. F.; Bonanni, L. J.; Clayburn, A. J.; Khera, Z.; Sadler, E. C.; Prasad, J.; Schlacter, J.; Liu, K.; Silva, B.; Montgomery, S.; Kim, E. J.; Lester, J.; Hill, T. M.; Avoricani, A.; Chervonski, E.; Davydov, J.; Small, W.; Chakravartty, E.; Grover, H.; Dodson, J. A.; Brody, A. A.; Aphinyanaphongs, Y.; Razavian, N.
Show abstract
ImportanceLarge language models (LLMs) are crucial for medical tasks. Ensuring their reliability is vital to avoid false results. Our study assesses two state-of-the-art LLMs (ChatGPT and LlaMA-2) for extracting clinical information, focusing on cognitive tests like MMSE and CDR. ObjectiveEvaluate ChatGPT and LlaMA-2 performance in extracting MMSE and CDR scores, including their associated dates. MethodsOur data consisted of 135,307 clinical notes (Jan 12th, 2010 to May 24th, 2023) mentioning MMSE, CDR, or MoCA. After applying inclusion criteria 34,465 notes remained, of which 765 underwent ChatGPT (GPT-4) and LlaMA-2, and 22 experts reviewed the responses. ChatGPT successfully extracted MMSE and CDR instances with dates from 742 notes. We used 20 notes for fine-tuning and training the reviewers. The remaining 722 were assigned to reviewers, with 309 each assigned to two reviewers simultaneously. Inter-rater-agreement (Fleiss Kappa), precision, recall, true/false negative rates, and accuracy were calculated. Our study follows TRIPOD reporting guidelines for model validation. ResultsFor MMSE information extraction, ChatGPT (vs. LlaMA-2) achieved accuracy of 83% (vs. 66.4%), sensitivity of 89.7% (vs. 69.9%), true-negative rates of 96% (vs 60.0%), and precision of 82.7% (vs 62.2%). For CDR the results were lower overall, with accuracy of 87.1% (vs. 74.5%), sensitivity of 84.3% (vs. 39.7%), true-negative rates of 99.8% (98.4%), and precision of 48.3% (vs. 16.1%). We qualitatively evaluated the MMSE errors of ChatGPT and LlaMA-2 on double-reviewed notes. LlaMA-2 errors included 27 cases of total hallucination, 19 cases of reporting other scores instead of MMSE, 25 missed scores, and 23 cases of reporting only the wrong date. In comparison, ChatGPTs errors included only 3 cases of total hallucination, 17 cases of wrong test reported instead of MMSE, and 19 cases of reporting a wrong date. ConclusionsIn this diagnostic/prognostic study of ChatGPT and LlaMA-2 for extracting cognitive exam dates and scores from clinical notes, ChatGPT exhibited high accuracy, with better performance compared to LlaMA-2. The use of LLMs could benefit dementia research and clinical care, by identifying eligible patients for treatments initialization or clinical trial enrollments. Rigorous evaluation of LLMs is crucial to understanding their capabilities and limitations.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 96%
- Quantifying Device Type and Handedness Biases in a Remote Parkinson’s Disease AI-Powered Assessment 93%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 93%
Similar papers in this journal
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 94%
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 94%
- Characterizing subgroup performance of probabilistic phenotype algorithms within older adults: A case study for dementia, mild cognitive impairment, and Alzheimer’s and Parkinson’s diseases 93%
Similar papers in this journal
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 93%
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 93%
- Design and implementation of a system for automated monitoring of adherence to evidenced-based clinical guideline recommendations 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.