ChatGPT for automating lung cancer staging: feasibility study on open radiology report dataset
Nakamura, Y.; Kikuchi, T.; Yamagishi, Y.; Hanaoka, S.; Nakao, T.; Miki, S.; Yoshikawa, T.; Abe, O.
Show abstract
ObjectivesCT imaging is essential in the initial staging of lung cancer. However, free-text radiology reports do not always directly mention clinical TNM stages. We explored the capability of OpenAIs ChatGPT to automate lung cancer staging from CT radiology reports. MethodsWe used MedTxt-RR-JA, a public de-identified dataset of 135 CT radiology reports for lung cancer. Two board-certified radiologists assigned clinical TNM stage for each radiology report by consensus. We used a part of the dataset to empirically determine the optimal prompt to guide ChatGPT. Using the remaining part of the dataset, we (i) compared the performance of two ChatGPT models (GPT-3.5 Turbo and GPT-4), (ii) compared the performance when the TNM classification rule was or was not presented in the prompt, and (iii) performed subgroup analysis regarding the T category. ResultsThe best accuracy scores were achieved by GPT-4 when it was presented with the TNM classification rule (52.2%, 78.9%, and 86.7% for the T, N, and M categories). Most ChatGPTs errors stemmed from challenges with numerical reasoning and insufficiency in anatomical or lexical knowledge. ConclusionsChatGPT has the potential to become a valuable tool for automating lung cancer staging. It can be a good practice to use GPT-4 and incorporate the TNM classification rule into the prompt. Future improvement of ChatGPT would involve supporting numerical reasoning and complementing knowledge. Clinical relevance statementChatGPTs performance for automating cancer staging still has room for enhancement, but further improvement would be helpful for individual patient care and secondary information usage for research purposes. Key pointsO_LIChatGPT, especially GPT-4, has the potential to automatically assign clinical TNM stage of lung cancer based on CT radiology reports. C_LIO_LIIt was beneficial to present the TNM classification rule to ChatGPT to improve the performance. C_LIO_LIChatGPT would further benefit from supporting numerical reasoning or providing anatomical knowledge. C_LI Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=119 SRC="FIGDIR/small/23299107v1_ufig1.gif" ALT="Figure 1"> View larger version (24K): org.highwire.dtl.DTLVardef@15e92c2org.highwire.dtl.DTLVardef@1f53392org.highwire.dtl.DTLVardef@10cf0b1org.highwire.dtl.DTLVardef@8e0150_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Predicting EGFR mutation status in lung adenocarcinoma presenting as ground-glass opacity: utilizing radiomics model in clinical translation 94%
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 94%
- From Community Acquired Pneumonia to COVID-19: A Deep Learning Based Method for Quantitative Analysis of COVID-19 on thick-section CT Scans 93%
Similar papers in this journal
- ai-corona : Radiologist-Assistant Deep Learning Framework for COVID-19 Diagnosis in Chest CT Scans 97%
- Weakly supervised learning for multi-organ adenocarcinoma classification in whole slide images 95%
- Enhancing Semantic Segmentation in Chest X-Ray Images through Image Preprocessing: ps-KDE for Pixel-wise Substitution by Kernel Density Estimation 94%
Similar papers in this journal
- Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning models 94%
- Classifying Non-Small Cell Lung Cancer Histopathology Types and Transcriptomic Subtypes using Convolutional Neural Networks 93%
- Natural language inference for clinical registry curation 92%
Similar papers in this journal
- Toward Understanding COVID-19 Pneumonia: A Deep-learning-based Approach for Severity Analysis and Monitoring the Disease 95%
- Effective Deep Learning Approaches for Predicting COVID-19 Outcomes from Chest Computed Tomography Volumes 94%
- COVID-Classifier: An automated machine learning model to assist in the diagnosis of COVID-19 infection in chest x-ray images 94%
Similar papers in this journal
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 96%
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 95%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.