Transforming Retinal Vascular Disease Classification: A Comprehensive Analysis of ChatGPT's Performance and Inference Abilities on Non-English Clinical Environment
Liu, X.; Wu, J.; Shao, A.; Shen, W.; Ye, P.; Wang, Y.; Ye, J.; Jin, K.; Yang, J.
Show abstract
ObjectiveTo evaluate the effectiveness and reasoning ability of ChatGPT in diagnosing retinal vascular diseases in the Chinese clinical environment. Materials and MethodsWe collected 1226 fundus fluorescein angiography reports and corresponding diagnosis written in Chinese, and tested ChatGPT with four prompting strategies (direct diagnosis or diagnosis with explanation and in Chinese or English). ResultsChatGPT using English prompt for direct diagnosis achieved the best performance, with F1-score of 80.05%, which was inferior to ophthalmologists (89.35%) but close to ophthalmologist interns (82.69%). Although ChatGPT can derive reasoning process with a low error rate, mistakes such as misinformation (1.96%), and hallucination (0.59%) still exist. Discussion and ConclusionsChatGPT can serve as a helpful medical assistant to provide diagnosis under non-English clinical environments, but there are still performance gaps, language disparity, and errors compared to professionals, which demonstrates the potential limitations and the desiration to continually explore more robust LLMs in ophthalmology practice.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Glaucoma Detection and Staging from Visual Field Images using Machine Learning Techniques 95%
- Diagnosis of central serous chorioretinopathy by deep learning analysis of en face images of choroidal vasculature 94%
- An automatic glaucoma grading method based on attention mechanism and EfficientNetB3 network 94%
Similar papers in this journal
- Unveiling the Clinical Incapabilities: A Benchmarking Study of GPT-4V(ision) for Ophthalmic Multimodal Image Analysis 97%
- Performance of DeepSeek-R1 in Ophthalmology: An Evaluation of Clinical Decision-Making and Cost-Effectiveness 95%
- AI-Powered Effective Lens Position Prediction Improves the Accuracy of Existing Lens Formulas 94%
Similar papers in this journal
- Automatic Measurements of Smooth Pursuit Eye Movements by Video-Oculography and Deep Learning-Based Object Detection 95%
- Low-cost, Smartphone-based Specular Imaging and Automated Analysis of the Corneal Endothelium 93%
- AutoMorph: Automated Retinal Vascular Morphology Quantification via a Deep Learning Pipeline 93%
Similar papers in this journal
- Association between Renal Function and Retinal Neurodegeneration in Chinese Patients with Type 2 Diabetes Mellitus 92%
- A Novel Triage Tool of Artificial Intelligence-Assisted Diagnosis Aid System for Suspected COVID-19 Pneumonia in Fever Clinics 90%
- Application of Telemedicine During the Coronavirus Disease Epidemics: A Rapid Review and Meta-Analysis 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.