Clinical Usability of Generative Artificial Intelligence for MR Safety Advice
Rose, H. E. L.; Thorpe, J. C.; Panek, R.; Goncalves, E.; Morgan, P. S.
Show abstract
This study investigated whether readily available, generative AI models, could be used to answer MR safety queries as an MR Safety Expert (MRSE), with "clinical usability" assessed by an expert review panel. This study is a mixed retrospective-prospective, proof-of-concept study. A clinical MR safety advice archive (January 2024 to April 2025) was used to curate 30 generic MR safety support requests with associated MRSE responses. ChatGPT-4o (ChatGPT) and Google AI Overview (GAIO) were prompted with these generic requests to generate AI safety advice. An expert panel assessed all answers for clinical usability. Unusable responses were assigned as; "Unsafe Advice", "Safe but Incorrect", "Incomplete Advice/ Key Details Missing", "Contradictory Statements"," Out of Date". Requests were subcategorised into "specific" and "generic" requests, as well as "passive" and "active" implants, and "other" requests for post review analysis. Percentages of usable answers and reasons for non-usable responses were compared. Overall, 93% (28/30) of the human responses, 50% (15/30) of the GAIO responses and 43% (13/30) of the ChatGPT responses were deemed acceptable for clinical use. Subcategorization usability was: "generic"; Human 94% (16/17), GAIO and ChatGPT 59% (10/17), "specific": Human 92% (12/13), GAIO 38% (5/13), ChatGPT 23% (3/13), Active: Human 100% (9/9), GAIO 33% (3/9), ChatGPT 22% (2/9), "passive": Human 88% (14/16), GAIO 56% (9/16), ChatGPT 50% (8/16) and "other"; Human 100% (5/5), GAIO and ChatGPT 60% (3/5). While both AIs were able provide clinically acceptable answers for some requests they did so at a significantly lower success rate than a human MRSE.
Matching journals
The top 11 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating Large Language Model-Generated Brain MRI Protocols: Performance of GPT4o, o3-mini, DeepSeek-R1 and Qwen2.5-72B 95%
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 91%
- First-generation clinical dual-source photon-counting CT: ultra-low dose quantitative spectral imaging 91%
Similar papers in this journal
- Inconsistency of AI in Intracranial Aneurysm Detection with Varying Dose and Image Reconstruction 94%
- MyoVision-US: an Artificial Intelligence-Powered Software for Automated Analysis of Skeletal Muscle Ultrasonography 91%
- Does contrast-enhancement improve visualisation of lenticulostriate arteries in cerebral small vessel disease using time-of-flight magnetic resonance angiography at 7 Tesla? 91%
Similar papers in this journal
- Large language model-based information extraction from free-text radiology reports: a scoping review protocol 89%
- Validity of intraoperative imageless navigation (Naviswiss™) for component positioning accuracy in primary total hip arthroplasty: Protocol for a prospective observational cohort study in a single-surgeon practice 89%
- Protocol of the observational study STRATUM-OS: First step in the development and validation of the STRATUM tool based on multimodal data processing to assist surgery in patients affected by intra-axial brain tumours 88%
Similar papers in this journal
- Necessity and Impact of Specialization of Large Foundation Model for Medical Segmentation Tasks 92%
- Quality Assurance Assessment of Intra-Acquisition Diffusion-Weighted and T2-Weighted Magnetic Resonance Imaging Registration and Contour Propagation for Head and Neck Cancer Radiotherapy 92%
- PSMA-Hornet: fully-automated, multi-target segmentation of healthy organs in PSMA PET/CT images 91%
Similar papers in this journal
- pyKNEEr: An image analysis workflow for open and reproducible research on femoral knee cartilage 92%
- Internal calibration for opportunistic computed tomography muscle density analysis 92%
- Increased brain coverage and efficiency when measuring current-induced magnetic fields by use of simultaneous multi-slice echo-planar MRI 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.