GPT-4, an artificial intelligence large language model,exhibits high levels of accuracy on dermatology specialty certificate exam questions.
Shetty, M.; Ettlinger, M.; Lynch, M.
Show abstract
Artificial Intelligence (AI) has shown considerable potential within medical fields including dermatology. In recent years a new form of AI, large language models, has shown impressive performance in complex textual reasoning across a wide range of domains including standardised medical licensing exam questions. Here, we compare the performance of different models within the GPT family (GPT-3, GPT-3.5, and GPT-4) on 89 publicly available sample questions from the Dermatology specialty certificate examination. We find that despite no specific training on dermatological text, GPT-4, the most advanced large language model, exhibits remarkable accuracy - answering in excess of 85% of questions correctly, at a level that would likely be sufficient to pass the SCE exam.
Matching journals
The top 12 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- MultiGML: Multimodal Graph Machine Learning for Prediction of Adverse Drug Events 89%
- Cloud-controlled microscopy enables remote project-based biology education in Latinx communities in the United States and Latin America 86%
- Comprehensive analysis of computational approaches in plant transcription factors binding regions discovery 86%
Similar papers in this journal
Similar papers in this journal
- Evaluating Anti-LGBTQIA+ Medical Bias in Large Language Models 92%
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 92%
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.