Artificial Intelligence in Neuro-Oncology: Assessing ChatGPT Accuracy in MRI Interpretation and Treatment Advice
Ishaque, A. H.; Boutet, A.; Hiremath, S. B.; Mullarkey, M. P.; Peris-Celda, M.; Zadeh, G.
Show abstract
PurposeLarge language models (LLMs) have demonstrated advanced capabilities in interpreting text and visual inputs. Their potential to transform oncological practice is significant, but their accuracy and reliability in interpreting medical imaging and offering management suggestions remain underexplored. This study aimed to evaluate the performance of ChatGPT in interpreting T1-weighted contrast-enhanced MRI images of meningiomas and glioblastomas and providing treatment recommendations based on simulated patient inquiries. MethodsThis observational cohort study utilized publicly available MRI datasets. Thirty cases of meningiomas and glioblastomas were randomly selected, yielding 90 images (three orthogonal planes per case). ChatGPT-4o was tasked with interpreting these images and responding to six standardized patient-simulated questions. Two neuroradiologists and neurosurgeons assessed ChatGPTs performance using five-point Likert scales and their inter-rater agreement was evaluated. ResultsChatGPT identified MRI sequences with 91.7% accuracy and localized tumors correctly in 66.7% of cases. Tumor size was qualitatively described in 85% of cases, and the median acceptability was rated as 4.0 (IQR 4.0-5.0) by neuroradiologists. ChatGPT included meningioma in the differential diagnosis for 73.3% of meningioma cases and glioma in 83.3% of glioblastoma cases. Inter-rater agreement among neuroradiologists ranged from moderate to good ({kappa} = 0.45-0.72). While surgical treatment was suggested in all symptomatic cases, neurosurgeon acceptability ratings varied, with poor inter-rater reliability. ConclusionsChatGPT demonstrates potential in interpreting neuro-oncological MRI images and offering preliminary management recommendations. However, errors in tumor localization and variability in recommendation acceptability underscore the need for physician oversight and further refinement of LLMs before clinical integration.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluation of an artificial intelligence model for detection of pneumothorax and tension pneumothorax on chest radiograph 90%
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 88%
- Characterizing Potential Conflicts of Interest Among UpToDate and DynaMed Content Contributors 88%
Similar papers in this journal
- Association of Extent of Resection and Functional Outcomes in Diffuse Low-Grade Glioma: Systematic Review & Meta-Analysis 91%
- Early Imaging Marker of Progressive Glioblastoma: a window of opportunity 90%
- Rapid DNA methylation-based classification of pediatric brain tumours from ultrasonic aspirate specimens 88%
Similar papers in this journal
- Automated Tumor Segmentation and Brain Tissue Extraction from Multiparametric MRI of Pediatric Brain Tumors: A Multi-Institutional Study 94%
- Early prognostication of overall survival for pediatric diffuse midline gliomas using MRI radiomics and machine learning 94%
- A New Method for Optimal Placement of Tumor Treating Fields Electrodes 93%
Similar papers in this journal
- Protocol of the observational study STRATUM-OS: First step in the development and validation of the STRATUM tool based on multimodal data processing to assist surgery in patients affected by intra-axial brain tumours 93%
- Protocol for the Tessa Jowell BRAIN MATRIX Platform Study 90%
- Large language model-based information extraction from free-text radiology reports: a scoping review protocol 90%
Similar papers in this journal
- Deep neural networks allow expert-level brain meningioma detection, segmentation and improvement of current clinical practice 94%
- Content-based image retrieval assists radiologists in diagnosing eye and orbital mass lesions in MRI 93%
- Comparison of Radiomic Feature Aggregation Methods for Patients with Multiple Tumors 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.