Evaluation of Large Language Models in the Clinical Management of Patients With Upper Gastrointestinal Bleeding: Insights From Real-World Patient Data
Rajabnia, M.; Davoodi, F.; Hajisafarali, M.; Asadinia, S.; Shiri, M.; Shirzad, F.; Saeedi, P.; Mohammadi, M.
Show abstract
ObjectiveUpper gastrointestinal bleeding (UGIB) is a life-threatening emergency requiring rapid risk assessment. Current scoring tools have limited accuracy. Large language models (LLMs) may support clinical decision-making, but their role in UGIB management is unclear. This study evaluated LLMs for patient risk classification, prediction of endoscopic findings, and alignment with routine clinical decision-making. MethodsIn this retrospective study, we analyzed electronic health records(EHRs) of 384 UGIB patients presented to two referral centers in Karaj, Iran, between March and December 2024. Included cases underwent upper gastrointestinal endoscopy; incomplete records were excluded. Five LLMs including GPT-5, Llama 4, Gemini-2.5-Flash, DeepSeek R1, and Grok were assessed using in-context learning for (i) risk classification, (ii) prediction of probable endoscopic findings, and (iii) clinical justification generation. Performance metrics included accuracy, precision, recall, and F1-score, compared with conventional machine learning models. Two gastroenterologists independently assessed justifications across seven domains: relevance, clarity, originality, completeness, specificity, correctness, and consistency. ResultsAll LLMs outperformed conventional models (highest baseline accuracy 0.54). GPT-5 achieved the highest risk classification accuracy (0.66), followed by Llama 4 (0.64). Grok performed best in predicting endoscopic findings (0.32). gastroenterologists noted variability in reasoning: GPT-5 and Grok provided the most complete justifications, though GPT-5 occasionally over-classified urgent cases. Llama-4 and Gemini-2.5-Flash were less specific, while DeepSeek R1 offered detailed patient summaries but lacked predictive outputs. ConclusionsLLMs improved UGIB risk prediction and generate interpretive reasoning, but accuracy limitations, inconsistent reasoning, and occasional risk misclassification highlight the need for clinician oversight and prospective validation before clinical use. Key MessagesO_ST_ABSWhat is already known on this topicC_ST_ABSUGIB is a medical emergency requiring rapid risk stratification and timely management. LLMs are promising tools for clinical decision support, but their role in UGIB management remains unclear. What this study addsLLMs can improve risk prediction and interpretive reasoning in UGIB, but limitations in accuracy, inconsistent reasoning, and occasional misclassification highlight the need for clinician oversight and prospective validation. How this study might affect research, practice, or policyLLMs provide structured, human-readable explanations that could support clinical decision-making, potentially reduce unnecessary emergency endoscopies, improve care efficiency, and alleviate physician workload.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Artificial Intelligence Model for Analyzing Colonic Endoscopy Images to Detect Changes Associated with Irritable Bowel Syndrome 94%
- Harnessing the Open Access Version of ChatGPT for Enhanced Clinical Opinions 92%
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 92%
Similar papers in this journal
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 93%
- A comparison of machine learning models versus clinical evaluation for mortality prediction in patients with sepsis 91%
- A method for rapid machine learning development for data mining with Doctor-In-The-Loop 91%
Similar papers in this journal
Similar papers in this journal
- Non-endoscopic screening for Barrett’s esophagus and Esophageal Adenocarcinoma in at risk Veterans 91%
- Nonendoscopic Detection Of Barrett’S Esophagus In Patients Without Gerd Symptoms 89%
- Prune intake ameliorates chronic constipation symptoms and causes little discomfort from diarrhea and loose stools: A randomized placebo-controlled trial 88%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.