Large Language Models Fail to Reproduce Level I Recommendations for Breast Radiotherapy
Tang, K.; Han, J.; Wu, S.
Show abstract
This study evaluates the reliability of the largest public-facing large language models in providing accurate breast cancer radiotherapy recommendations. We assessed ChatGPT 3.5, ChatGPT 4, ChatGPT 4o, Claude 3.5 Sonnet, and ChatGPT o1 in three common clinical scenarios. The clinical cases are as follows: post-lumpectomy radiotherapy in a 40 year old woman, (2) postmastectomy radiation in a 40 year old woman with 4+ lymph nodes, and (3) postmastectomy radiation in an 80 year old woman with early stage tumor and negative axillary dissection. Each case was designed to be unambiguous with respect to the Level I evidence and clinical guideline-supported approach. The evidence-supported radiation treatments are as follows: (1) Whole breast with boost (2) Regional nodal irradiation (3) Omission of post-operative radiotherapy. Each prompt is presented to each LLM multiple times to ensure reproducibility. Results indicate that the free, public-facing models often fail to provide accurate treatment recommendations, particularly when omission of radiotherapy was the correct course of action. Many recommendations suggested by the LLMs increase morbidity and mortality in patients. Models only accessible through paid subscription (ChatGPT o1 and o1-mini) demonstrated greatly improved accuracy. Some prompt-engineering techniques, rewording and chain-of-reasoning, enhanced the accuracy of the LLMs, while true/false questioning significantly worsened results. While public-facing LLMs show potential for medical applications, their current reliability is unsuitable for clinical decision-making.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Comprehensive Quantitative Evaluation of Inter-observer Delineation Performance of MR-guided Delineation of Oropharyngeal Gross Tumor Volumes and High-risk Clinical Target Therapy: An R-IDEAL Stage 0 Prospective Study 96%
- Normal Tissue Complication Probability (NTCP) prediction model for osteoradionecrosis of the mandible in head and neck cancer patients following radiotherapy: Large-scale observational cohort 95%
- Clinical Impact of Contouring Variability for Prostate Cancer Tumor Boost 95%
Similar papers in this journal
Similar papers in this journal
- Predictive factors for the development of peritumoral brain edema after LINAC-based radiation treatment in patients with intracranial meningioma 93%
- National diagnostic reference levels for digital diagnostic and screening mammography in Uganda. 93%
- Classification performance bias between training and test sets in a limited mammography dataset 92%
Similar papers in this journal
- Dosimetric Analysis of Fast Forward Breast Radiotherapy Using 3D-CRT with Deep Inspiration Breath Hold(DIBH) 94%
- Benchmarking Deep Learning-based Image Retrieval of Oral Tumor Histology 90%
- SARS-CoV-2 antibody seroprevalence in cancer patients on systemic antineoplastic treatment in the first wave of the COVID-19 pandemic in Portugal 89%
Similar papers in this journal
- Artificial Intelligence Uncertainty Quantification in Radiotherapy Applications - A Scoping Review 96%
- Detailed patient-individual reporting of lymph node involvement in oropharyngeal squamous cell carcinoma with an online interface 94%
- Evaluation of indirect damage and damage saturation effects in dose-response curves of hypofractionated radiotherapy of early-stage NSCLC and brain metastases 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.