Evaluating the Diagnostic and Treatment Recommendation Capabilities of GPT-4 Vision in Dermatology
Pillai, A.; Parappally-Joseph, S.; Hardin, J.
Show abstract
BackgroundThe integration of artificial intelligence (AI) in dermatology presents a promising frontier for enhancing diagnostic accuracy and treatment planning. However, general purpose AI models require rigorous evaluation before being applied to real-world medical cases. ObjectiveThis project specifically evaluates GPT-4Vs performance in accurately diagnosing and generating treatment plans for common dermatological conditions, comparing its assessment of textual versus image data and its performance with multimodal inputs. Beyond the immediate scope, this study contributes to the broader trajectory of integrating AI in healthcare, highlighting the limitations of these technologies, as well as their potential to enhance efficiency, and education within medical training and practice. MethodsA dataset of 102 images representing nine common dermatological conditions was compiled from open-access websites. Fifty-four images were ultimately selected by two board-certified dermatologists as being representative and typical of the common conditions. Additionally, nine clinical scenarios corresponding to these conditions were developed. GPT-4Vs diagnostic capabilities were assessed in three setups: Image Prompt (image-based), Scenario Prompt (text-based), and Image and Scenario Prompt (combining both modalities). The models performance was evaluated based on diagnostic accuracy, differential diagnosis, and treatment recommendations. ResultsIn the Image Prompt setup, GPT-4V correctly identified the primary diagnosis for 29 of 54 images. The Scenario Prompt setup showed a higher accuracy rate of 89% in identifying the primary diagnosis. The multimodal Image and Scenario Prompt setup also achieved an 89% accuracy rate. However, a notable bias towards textual data over visual data was observed. Treatment recommendations were evaluated by the same two dermatologists, using a modified Entrustment Scale, showing competent but not expert-level performance. ConclusionGPT-4V demonstrates promising capabilities in dermatological diagnosis and treatment recommendations, particularly in text-based scenarios. However, its performance in image-based diagnosis and integration of multimodal data highlights areas for improvement. The study underscores the potential of AI in augmenting dermatological practice, emphasizing the need for further development, and fine-tuning of such models to ensure their efficacy and reliability in clinical settings.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The mathematics of erythema: Development of machine learning models for artificial intelligence assisted measurement and severity scoring of radiation induced dermatitis 92%
- Refining LLMs Outputs with Iterative Consensus Ensemble (ICE) 92%
- SymScore: Machine Learning Accuracy Meets Transparency in a Symbolic Regression-Based Clinical Score Generator 89%
Similar papers in this journal
- Evaluation of a clinical decision support system for detection of patients at risk after kidney transplantation 90%
- Applying machine-learning to rapidly analyse large qualitative text datasets to inform the COVID-19 pandemic response: Comparing human and machine-assisted topic analysis techniques 89%
- AI-Driven Early Detection of Severe Influenza in Jiangsu, China: A Deep Learning Model Validated Through The Design of Multi-Center Clinical Trials and Prospective Real-World Deployment 88%
Similar papers in this journal
- Computer vision detects inflammatory arthritis in standardized smartphone photographs in an Indian patient cohort 89%
- Development and External Validation of a Mixed-Effects Deep Learning Model to Diagnose COVID-19 from CT Imaging 89%
- Emerging Applications of NLP and Large Language Models in Gastroenterology and Hepatology: A Systematic Review 88%
Similar papers in this journal
Similar papers in this journal
- An Inherently Interpretable AI model improves Screening Speed and Accuracy for Early Diabetic Retinopathy 92%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 92%
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.