Close, But no Cigar: Comparative Evaluation of ChatGPT-4o and OpenAI o1-preview in Answering Pancreatic Ductal Adenocarcinoma-Related Questions
Li, C.; Chu, Y.; Liu, D.; Ghanad, E.; Abdelhadi, S.; Sandra-Petrescu, F.; Reissfelder, C.; Vassilev, G.; Yang, C.
Show abstract
BackgroundThis study aimed to evaluate the effectiveness of ChatGPT-4o and OpenAI o1-preview in responding to pancreatic ductal adenocarcinoma (PDAC)-related queries. The study assessed both LLMs accuracy, comprehensiveness, and safety when answering clinical questions, based on the National Comprehensive Cancer Network(R) (NCCN) Clinical Practice Guidelines for PDAC. MethodsThe study used a 20-question dataset derived from clinical scenarios related to PDAC. Two board-certified surgeons independently evaluated the responses by ChatGPT-4o and OpenAI o1-preview for their accuracy, comprehensiveness, and safety using a Likert scale. Statistical analyses were conducted to compare the performances of the two models. We also analyzed the impact of OpenAI o1-previews Chain of Thought (CoT) technology. ResultsBoth models demonstrated high median scores across all dimensions (5 out of 5). OpenAI o1-preview outperformed ChatGPT-4o in comprehensiveness (p = 0.026) and demonstrated superior reasoning ability, with a higher accuracy rate of 75% compared to 60% for ChatGPT-4o. OpenAI o1-preview generated more concise responses (median 64 vs. 82 words, p < 0.001). The CoT method in OpenAI o1-preview appeared to enhance its reasoning capabilities, particularly in complex treatment decisions. However, both models made critical errors in some complex clinical scenarios. ConclusionOpenAI o1-preview, with its CoT technology, demonstrates higher comprehensiveness than ChatGPT-4.0 and showed a tendency of improved accuracy. However, both models still make critical errors and cause some harm to patients. Even the most advanced models are not suitable for offering reliable medical information and cannot function as an assistant for decision-making.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Can Large Language Models Aid Caregivers of Pediatric Cancer Patients in Information Seeking? A Cross-Sectional Investigation 95%
- A Video Intervention to Improve Patient Understanding of Tumor Genomic Testing in Patients with Cancer 90%
- Making prognostic algorithms useful in shared decision-making: Patients and clinicians’ requirements for the Predict:Breast Cancer interface 89%
Similar papers in this journal
- Predicting Axillary Lymph Node Metastasis in Early Breast Cancer Using Deep Learning on Primary Tumor Biopsy Slides 91%
- Diagnostic accuracy and safety of coaxial core-needle biopsy (CNB) system in Oncology patients treated in a specialist cancer centre with prospective validation within clinical trial data 90%
- Factors affecting COVID-19 outcomes in cancer patients - A first report from Guys Cancer Centre in London 90%
Similar papers in this journal
- The NILS study protocol - a retrospective validation study of a preoperative decision-making tool for non-invasive lymph node staging in women with primary breast cancer [ISRCTN14341750] 90%
- Demarcation line determination for diagnosis of gastric cancer disease range using unsupervised machine learning in magnifying narrow-band imaging 89%
- Deep learning models for poorly differentiated colorectal adenocarcinoma classification in whole slide images using transfer learning 89%
Similar papers in this journal
- The role of natural language processing in cancer care: a systematic scoping review with narrative synthesis 95%
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 93%
- Machine learning in the identification of prognostic DNA methylation biomarkers among patients with cancer: a systematic review of epigenome-wide studies 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.