Expert-Guided Visual Correction for Characterizing Diagnostic Performance and Error Patterns of Multimodal Large Language Models Using Periodontal In-Service Examination Images
Dhaimade, P. A.; Henderson, R.
Show abstract
Multimodal large language models (MLLMs) are increasingly applied to image-based clinical reasoning, yet their diagnostic reliability in periodontal image interpretation, and the underlying source of their errors, remain poorly characterized. This study evaluated six architecturally distinct MLLMs (Claude Sonnet 4.5, GPT-5.0, Gemini 2.5, GLM-4.6, Sonar, and Grok 4.1) using 50 image-based multiple-choice questions drawn from the American Academy of Periodontology In-Service Examination, spanning clinical photographs, histopathology, radiographs, cardiac rhythm strips, and anatomical illustrations. A sequential two-phase experimental design was used: in Phase 1, each model independently described each image, selected an answer, and provided a supporting citation; in Phase 2, applied only to questions answered incorrectly, models were given an expert-validated visual description and asked to re-answer, allowing diagnostic improvement through visual correction to be measured directly. Expert ground truth for image content was established by a board-certified periodontist and independently validated by a second board-certified periodontist. Model outputs were classified using a dual-process error taxonomy adapted from Norman's model of diagnostic reasoning, distinguishing perceptual errors, arising from inaccurate visual feature extraction, from cognitive errors, arising from flawed reasoning despite accurate perception, with cognitive errors further subdivided into correctable and persistent subtypes, and additional categories capturing compound perceptual-cognitive failures and compensatory reasoning that overcame inaccurate perception. Diagnostic accuracy and error type distribution varied significantly across models and image modality. Correcting inaccurate visual descriptions in Phase 2 improved diagnostic accuracy for a subset of previously incorrect responses, indicating that a meaningful share of errors originated at the level of visual perception rather than clinical reasoning; conversely, a distinct subset of errors persisted despite accurate corrected visual input, indicating reasoning-level failures independent of perceptual accuracy. Some models also reached correct answers despite generating inaccurate image descriptions, reflecting compensatory reasoning resilient to perceptual error. These findings show that aggregate accuracy scores conflate mechanistically distinct failure modes, and that perceptual and cognitive errors carry different implications for how MLLMs might be safely deployed or improved for diagnostic image interpretation. The expert-guided visual correction framework introduced here provides a generalizable, mechanism-based approach to benchmarking multimodal AI diagnostic performance that extends beyond periodontics to other visually driven diagnostic domains in medicine. As MLLMs become increasingly accessible to clinicians, residents, and dental educators, distinguishing perceptual from cognitive failure is essential for guiding responsible clinical use, targeting model refinement, and informing AI-augmented dental education and competency assessment.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 90%
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 90%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 90%
Similar papers in this journal
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 87%
- A Crowdsourcing Approach to Develop Machine Learning Models to Quantify Radiographic Joint Damage in Rheumatoid Arthritis 87%
- Diagnostic Codes in AI prediction models and Label Leakage of Same-admission Clinical Outcomes 87%
Similar papers in this journal
Similar papers in this journal
- Utilization of Generative AI-drafted Responses for Managing Patient-Provider Communication 91%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 90%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.