GPT-4V(ision) Unsuitable for Clinical Care and Education: A Clinician-Evaluated Assessment
Senkaiahliyan, S.; Toma, A.; Ma, J.; Chan, A.-W.; Ha, A.; An, K. R.; Suresh, H.; Rubin, B.; Wang, B.
Show abstract
OpenAIs large multimodal model, GPT-4V(ision), was recently developed for general image interpretation. However, less is known about its capabilities with medical image interpretation and diagnosis. Board-certified physicians and senior residents assessed GPT-4Vs proficiency across a range of medical conditions using imaging modalities such as CT scans, MRIs, ECGs, and clinical photographs. Although GPT-4V is able to identify and explain medical images, its diagnostic accuracy and clinical decision-making abilities are poor, posing risks to patient safety. Despite the potential that large language models may have in enhancing medical education and delivery, the current limitations of GPT-4V in interpreting medical images reinforces the importance of appropriate caution when using it for clinical decision-making.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 95%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 94%
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 94%
Similar papers in this journal
- Large language models for generating medical examinations: systematic review 95%
- Performance of ChatGPT on Chinese National Medical Licensing Examinations: A Five-Year Examination Evaluation Study for Physicians, Pharmacists and Nurses 93%
- Student self-assessment: feasibility, advantages and limitations Example of a workshop for trainee surgeons using a suture score 92%
Similar papers in this journal
- Multimodal Pain Recognition in Postoperative Patients: A Machine Learning Approach 92%
- The potential for digital patient symptom recording through symptom assessment applications to optimize patient flow and reduce waiting times in Urgent Care Centers: a simulation study 91%
- Improving emergency department patient-doctor conversation through an artificial intelligence symptom taking tool: an action-oriented design pilot study 91%
Similar papers in this journal
- ChatGPT Provides Inconsistent Risk-Stratification of Patients With Atraumatic Chest Pain 93%
- Towards implementation of AI in New Zealand national screening program: Cloud-based, Robust, and Bespoke 93%
- ai-corona : Radiologist-Assistant Deep Learning Framework for COVID-19 Diagnosis in Chest CT Scans 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.