A Systematic Performance Evaluation of Three Large Language Models in Answering Questions on moderate Hyperthermia
Dennstaedt, F.; Cihoric, N.; Bachmann, N.; Filchenko, I.; Berclaz, L.; Crezee, H.; Curto, S.; Ghadjar, P.; Huebenthal, B.; Hurwitz, M. D.; Kok, P.; Lindner, L. H.; Marder, D.; Molitoris, J.; Notter, M.; Rahman, S.; Riesterer, O.; Spalek, M.; Trefna, H.; Zilli, T.; Rodrigues, D.; Fuerstner, M.; Stutz, E.
Show abstract
BackgroundLarge Language Models (LLMs) have demonstrated expert-level performance across many medical domains, suggesting potential utility in clinical practice. However, their reliability in the highly specialized domain of moderate hyperthermia (HT) remains unknown. We therefore evaluated the performance of three modern LLMs in answering HT-related questions. MethodsWe conducted an evaluation study by posing 40 open-ended questions--22 clinical and 18 physics-related--to three modern LLMs (DeepSeek-V3, Llama-3.3-70B-Instruct, and GPT-4o). Responses were blinded, randomized, and evaluated by 19 international experts with either a clinical or physics background for quality (5-point Likert scale: 1=very bad, 2=bad, 3=acceptable, 4=good to 5=very good) and for potential harmfulness in clinical decision-making. ResultsA total of 1144 quality evaluation responses were collected. Overall reported mean quality scores were similar across models, with DeepSeek scoring 3.26, Llama 3.18, and GPT-4o 3.07, corresponding to an "acceptable" rating. Across expert evaluations, responses were considered potentially harmful in 17.8% of cases for DeepSeek, 19.3% for Llama, and 15.3% for GPT-4o. Notably, despite "acceptable" mean scores, approximately 25% of responses were rated "bad" to "very bad," and potentially harmful answers occurred in [~]15-19% of evaluations, indicating a non-trivial risk if used without domain expertise. ConclusionOur findings indicate that the performance of LLMs in HT in versions available at the time of investigation is only partially satisfactory. The proportion of poor-quality responses is too high and may lead non-domain experts to misinterpret the available clinical evidence and draw inappropriate clinical conclusions.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Cluster-Based Toxicity Estimation of Osteoradionecrosis via Unsupervised Machine Learning: Moving Beyond Single Dose-Parameter Normal Tissue Complication Probability by Using Whole Dose-Volume Histograms for Cohort Risk Stratification 95%
- Comprehensive Quantitative Evaluation of Inter-observer Delineation Performance of MR-guided Delineation of Oropharyngeal Gross Tumor Volumes and High-risk Clinical Target Therapy: An R-IDEAL Stage 0 Prospective Study 95%
- Normal Tissue Complication Probability (NTCP) prediction model for osteoradionecrosis of the mandible in head and neck cancer patients following radiotherapy: Large-scale observational cohort 94%
Similar papers in this journal
- Artificial Intelligence Uncertainty Quantification in Radiotherapy Applications - A Scoping Review 98%
- Serum miRNA-based signature indicates radiation exposure and dose in humans: a multicenter diagnostic biomarker study 93%
- LITE SABR M1: a Phase I Trial of Lattice Stereotactic Body Radiotherapy for Large Tumors 93%
Similar papers in this journal
- Development of a High-Performance Multiparametric MRI Oropharyngeal Primary Tumor Auto-Segmentation Deep Learning Model and Investigation of Input Channel Effects: Results from a Prospective Imaging Registry 94%
- Personalized volume-deescalated elective nodal irradiation in oropharyngeal squamous cell carcinoma (DeEscO): a study protocol 94%
- Morphological changes after cranial fractionated photon radiotherapy: localized loss of white matter and grey matter volume with increasing dose 94%
Similar papers in this journal
- Prediction of radiation-induced hypothyroidism using radiomic data analysis does not show superiority over standard normal tissue complication models 95%
- Standardising Breast Radiotherapy Structure Naming Conventions: A Machine Learning Approach 94%
- Quality of life and patient-reported outcomes following proton therapy for oropharyngeal carcinoma: a systematic review 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.