Artificial Intelligence in Periodontology: Performance Evaluation of ChatGPT, Claude, and Gemini on the In-service Examination
Ahmad, B.; Saleh, K.; Alharbi, S.; Alqaderi, H.; Jeong, Y. N.
Show abstract
BackgroundArtificial intelligence (AI) language models have shown potential as educational tools in healthcare, but their accuracy and reliability in periodontology education require further evaluation. In this study we aimed to assess and compare the performance of three prominent AI language models--ChatGPT-4o, Claude 3 Opus, and Gemini Advanced--with second-year periodontics residents across the United States on the American Academy of Periodontology 2024 in-service examination. MethodsWe conducted a cross-sectional study using 331 multiple-choice questions from the 2024 periodontology in-service examination. We evaluated and compared the performances of ChatGPT-4o, Claude 3 Opus, and Gemini Advanced across various question domains. The results of second-year periodontics residents served as a benchmark. ResultsChatGPT-4o, Gemini Advanced, and Claude 3 Opus significantly outperformed second-year periodontics residents across the United States, with accuracy rates of 92.7 percent, 81.6 percent, and 78.5 percent, respectively, compared to the residents 61.9 percent. The differences in performance among the AI models were statistically significant (p < 0.001). Percentile rankings underscored the superior performance of the AI models, with ChatGPT-4o, Gemini Advanced, and Claude 3 Opus placing in the 99.95th, 98th, and 95th percentiles, respectively. ConclusionChatGPT-4o displayed superior performance compared to Claude 3 Opus and Gemini Advanced. The results highlight the potential of AI large language models (LLMs) as educational tools in periodontology and emphasize the need for ongoing evaluation and validation as these technologies evolve. Researchers should explore both the integration of AI language models into periodontal education and their impact on learning outcomes and clinical decision-making.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 92%
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 91%
- Use of assistive technology to assess distal motor function in subjects with neuromuscular disease 90%
Similar papers in this journal
- Periodontal inflammation mediates the link between homocysteine and high blood pressure 88%
- Evaluating the potential of marine invertebrate and insect protein hydrolysates to reduce fetal bovine serum in cell culture media for cultivated fish production 85%
- SiCoDEA: a simple, fast and complete app for analyzing the effect of individual drugs and their combinations 85%
Similar papers in this journal
- Knowledge and aptitude of early childhood, primary and/or secondary education teachers referred to first aid measures in dental trauma in the province of Seville (Spain.) 93%
- Language comprehension developmental milestones in typically developing children assessed by the new Language Phenotype Assessment (LPA) 84%
- Different features of cholera in malnourished and non-malnourised children: analysis of 10-year surveillance data from a large diarrheal disease hospital in urban Bangladesh 84%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.