Zero-Shot Prompting is the Most Accurate and Scalable Strategy for Abstracting the Mayo Endoscopic Subscore from Colonoscopy Reports Using GPT-4
Yim, R. P.; Rudrapatna, V. A.
Show abstract
Structured AbstractO_ST_ABSIntroductionC_ST_ABSLarge-language models can help extract information from clinical notes, making them potentially useful for research in ulcerative colitis. However, it remains unclear if these models will scale well in practice. MethodsWe analyzed the performance and cost of programmatically using GPT-4 to abstract Mayo endoscopic subscores (MES) from 499 colonoscopy reports using different prompting strategies. ResultsZero-shot prompting, where GPT-4 is instructed without examples, was most accurate (83.55%) and cost-effective ($0.097/note). DiscussionUsing GPT-4 to automatically curate the MES and other variables is a practical strategy for quantifying UC activity and measuring improvements to clinical care.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 95%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 95%
- Design and Implementation of an End-to-End AI-Driven Colonoscopy Recall Workflow at Scale 94%
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 94%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.