Performance of GPT-4V(ision) in Ophthalmology: Use of Images in Clinical Questions
Tomita, K.; Nishida, T.; Kitaguchi, Y.; Miyake, M.; Kitazawa, K.
Show abstract
Background/aimsTo compare the diagnostic accuracy of Generative Pre-trained Transformer with Vision (GPT)-4 and GPT-4 with Vision (GPT-4V) for clinical questions in ophthalmology. MethodsThe questions were collected from the "Diagnosis This" section on the American Academy of Ophthalmology website. We tested 580 questions and presented GPT-4V with the same questions under two conditions: 1) multimodal model, incorporating both the question text and associated images, and 2) text-only model. We then compared the difference in accuracy between the two conditions using the chi-square test. The percentage of general correct answers was also collected from the website. ResultsThe GPT-4V model demonstrated higher accuracy with images (71.7%) than without images (66.7%, p<0.001). Both GPT-4 models showed higher accuracy than the general correct answers on the website [64.6 (95%CI, 62.9 to 66.3)]. ConclusionsThe addition of information from images enhances the performance of GPT-4V in diagnosing clinical questions in ophthalmology. This suggests that integrating multimodal data could be crucial in developing more effective and reliable diagnostic tools in medical fields. SYNOPSISThe study compared the diagnostic accuracy of GPT-4 and GPT-4 with Vision for clinical questions in ophthalmology, finding that the performance improved when it analyzed both text and images. WHAT IS ALREADY KNOWN ON THIS TOPICText-based large language models (LLMs) have demonstrated significant potential in enhancing medical interpretation and diagnosis. Generative Pretrained Transformer 4 with Vision (GPT-4V) can address image-related questions, but the use of GPT-4V in ophthalmology has not yet been validated. WHAT THIS STUDY ADDSOur study reports the answer accuracy on Diagnose This, provided by the American Academy of Ophthalmology, using GPT-4V. The integration of image data with GPT-4V enhances diagnostic accuracy in addressing ophthalmic clinical questions. HOW THIS STUDY MIGHT AFFECT RESEARCH, PRACTICE OR POLICYOur study indicates that combining image data with GPT-4 can enhance diagnostic accuracy in ophthalmic clinical questions. The development of LLMs trained on medical-specific datasets could further increase accuracy, advancing towards practical clinical applications.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Detecting papilloedema as a marker of raised intracranial pressure using artificial intelligence: a systematic review 95%
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 93%
- An Inherently Interpretable AI model improves Screening Speed and Accuracy for Early Diabetic Retinopathy 93%
Similar papers in this journal
- Unveiling the Clinical Incapabilities: A Benchmarking Study of GPT-4V(ision) for Ophthalmic Multimodal Image Analysis 98%
- Autonomous Screening for Laser Photocoagulation in Fundus Images Using Deep Learning 95%
- Performance of DeepSeek-R1 in Ophthalmology: An Evaluation of Clinical Decision-Making and Cost-Effectiveness 94%
Similar papers in this journal
- Towards implementation of AI in New Zealand national screening program: Cloud-based, Robust, and Bespoke 96%
- Glaucoma Detection and Staging from Visual Field Images using Machine Learning Techniques 95%
- Prediction of the ectasia screening index from raw Casia2 volume data for keratoconus identification by using convolutional neural networks 94%
Similar papers in this journal
- Automatic Measurements of Smooth Pursuit Eye Movements by Video-Oculography and Deep Learning-Based Object Detection 94%
- Low-cost, Smartphone-based Specular Imaging and Automated Analysis of the Corneal Endothelium 93%
- Visual field evaluation using Zippy Adaptive Threshold Algorithm (ZATA) Standard and ZATA Fast in patients with glaucoma and healthy individuals 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.