Evaluating GPT-4o Model Proficiency and Clinical Reasoning for Antimicrobial Stewardship in Dentistry
Dick, M.; Madathil, S.; Patel, A.; Kapoor, H. S.; Sharma, M.; D'Souza, Z.; Hameed, S.; Abu-Samak, M.; Najirad, A.; Dwairi, D.; Radaideh, O.; Nicolau, B.
Show abstract
Objectives: Dentists prescribe approximately one in ten antibiotics worldwide, yet antimicrobial stewardship (AMS) remains underemphasized in dental education. Large language models (LLMs) may support AMS training, but their proficiency and clinical reasoning in this context remain unclear. We evaluated GPT-4o's accuracy and clinical reasoning on dental antibiotic prescribing questions, stratified by question difficulty. Methods: We assembled 125 multiple-choice questions on dental antibiotic prescribing from eight peer-reviewed studies (2017-2023). GPT-4o answered each question and generated a clinical justification. Accuracy was assessed against source-study answer keys and examined across difficulty quartiles. Justifications were evaluated using an adapted 12-axis human-evaluation framework assessing scientific consensus, extent and likelihood of harm, inappropriate and missing content, bias, and both correct and incorrect comprehension, retrieval, and reasoning. Prophylaxis-specific questions were analysed separately. Results: GPT-4o correctly answered 72% of questions. Accuracy remained relatively stable across difficulty quartiles (78%, 78%, 65%, 70%). Experts rated 95.4% of justifications positively across the 12 axes. Comprehension, retrieval, and reasoning each exceeded 96.2% positive ratings. Missing content was the main weakness (7.8%), and 7.1% of justifications showed a moderate-to-severe potential for harm. Performance on prophylaxis-specific questions (98.1%) exceeded non-prophylaxis questions (93.0%). Conclusions: GPT-4o demonstrated moderate-to-high proficiency and clinically defensible reasoning in dental antibiotic prescribing questions. However, residual risks indicate that it is not suitable for unsupervised clinical use but shows potential as a supervised AMS educational tool.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Silicone toothbrushes: A scoping review of an underutilized tool in global oral health 90%
- A qualitative study on factors influencing health workers’ uptake of a pilot surgical antibiotic prophylaxis stewardship programme in selected Georgian hospitals 89%
- Advancing the Safe Motherhood Initiative: a qualitative and sentiment analysis of local physician’s perspectives on antibiotic self-medication during pregnancy in a low- and middle-income country 89%
Similar papers in this journal
- Knowledge and aptitude of early childhood, primary and/or secondary education teachers referred to first aid measures in dental trauma in the province of Seville (Spain.) 89%
- Language comprehension developmental milestones in typically developing children assessed by the new Language Phenotype Assessment (LPA) 86%
- Different features of cholera in malnourished and non-malnourised children: analysis of 10-year surveillance data from a large diarrheal disease hospital in urban Bangladesh 82%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.