Assessing Large Language Models for Oncology Data Inference from Radiology Reports
Chen, L.-C.; Zack, T.; Demirci, A.; Sushil, M.; Miao, B.; Kasap, C.; Butte, A. J.; Collisson, E.; Hong, J.
Show abstract
PurposeWe examined the effectiveness of proprietary and open Large Language Models (LLMs) in detecting disease presence, location, and treatment response in pancreatic cancer from radiology reports. MethodsWe analyzed 203 deidentified radiology reports, manually annotated for disease status, location, and indeterminate nodules needing follow-up. Utilizing GPT-4, GPT-3.5-turbo, and open models like Gemma-7B and Llama3-8B, we employed strategies such as ablation and prompt engineering to boost accuracy. Discrepancies between human and model interpretations were reviewed by a secondary oncologist. ResultsAmong 164 pancreatic adenocarcinoma patients, GPT-4 showed the highest accuracy in inferring disease status, achieving a 75.5% correctness (F1-micro). Open models Mistral-7B and Llama3-8B performed comparably, with accuracies of 68.6% and 61.4%, respectively. Mistral-7B excelled in deriving correct inferences from "Objective Findings" directly. Most tested models demonstrated proficiency in identifying disease containing anatomical locations from a list of choices, with GPT-4 and Llama3-8B showing near parity in precision and recall for disease site identification. However, open models struggled with differentiating benign from malignant post-surgical changes, impacting their precision in identifying findings indeterminate for cancer. A secondary review occasionally favored GPT-3.5s interpretations, indicating the variability in human judgment. ConclusionLLMs, especially GPT-4, are proficient in deriving oncological insights from radiology reports. Their performance is enhanced by effective summarization strategies, demonstrating their potential in clinical support and healthcare analytics. This study also underscores the possibility of zero-shot open model utility in environments where proprietary models are restricted. Finally, by providing a set of annotated radiology reports, this paper presents a valuable dataset for further LLM research in oncology.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DeepPhe-CR: Natural Language Processing Software Services for Cancer Registrar Case Abstraction 94%
- Large language models to help appeal denied radiotherapy services 93%
- Histology-based Prediction of Therapy Response to Neoadjuvant Chemotherapy for Esophageal and Esophagogastric Junction Adenocarcinomas Using Deep Learning 92%
Similar papers in this journal
- EchoGraph System for Automated Quality Assessment of Echocardiography Reports 94%
- A human-in-the-loop explanation framework for morphologically transparent AI predictions from whole-slide images 93%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 93%
Similar papers in this journal
- Large Language Models Improve the Identification of Emergency Department Visits for Symptomatic Kidney Stones 94%
- EHR Foundation Models Improve Robustness in the Presence of Temporal Distribution Shift 93%
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 93%
Similar papers in this journal
- A Crowdsourcing Approach to Develop Machine Learning Models to Quantify Radiographic Joint Damage in Rheumatoid Arthritis 90%
- Diagnostic Codes in AI prediction models and Label Leakage of Same-admission Clinical Outcomes 90%
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 90%
Similar papers in this journal
- Extraction of Crohn's Disease Clinical Phenotypes from Clinical Text Using Natural Language Processing 93%
- A user-friendly tool for cloud-based whole slide image segmentation, with examples from renal histopathology 92%
- Achieving Inclusive Healthcare through Integrating Education and Research with AI and Personalized Curricula 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.