Back

Automated Assessment of OSCE Physical Exams using Multimodal AI

Kang, S.; Holcomb, M.; Shakur, A. H.; Hein, D.; Ngo, H.-T.; Schuler, H.; Jarrett, P.; Dalton, T.; Jamieson, A. R.

2026-01-18 medical education
10.64898/2026.01.09.26343786 medRxiv
Show abstract

BackgroundThe assessment of physical examination skills in medical education is resource-intensive and prone to inter-rater variability. While artificial intelligence (AI) has successfully automated the grading of clinical notes and transcripts, evaluating the physical techniques themselves--what students do rather than what they say--remains an unsolved challenge. We evaluated whether a multimodal AI system could assess physical examination skills with expert-level reliability. MethodsIn this retrospective ablation study, we analyzed 300 video-recorded encounters from six Objective Structured Clinical Examination (OSCE) stations (cardiovascular, respiratory, gastrointestinal, musculoskeletal, and neurological). We compared the performance of a multimodal AI model (Gemini 2.5 Pro) across single- and multi-camera configurations and isolated input modalities (video, audio-only, transcript-only, visual-only) against standard human grading. The primary outcome was agreement with a physician-adjudicated ground-truth reference standard, measured by quadratic weighted Cohens kappa ({kappa}). ResultsThe AI system using a synchronized 3-camera native video configuration achieved significantly higher reliability ({kappa} = 0.830; 95% CI, 0.773-0.880) than the standard human evaluators ({kappa} = 0.732; 95% CI, 0.687-0.776). Performance followed a strict hierarchy: native video > audio-only > transcript-only > visual-only. Notably, visual-only models failed ({kappa} {approx} 0.20) despite high detection accuracy, revealing a "visual paradox" where models could identify when an action occurred but not how well it was performed without audio cues. ConclusionsA properly configured multimodal AI system can grade physical examination skills with reliability exceeding that of trained human evaluators. Success requires native processing of synchronized audio-visual streams; transcript-based or visual-only approaches are insufficient for high-stakes assessment. These findings suggest that AI can provide scalable, objective, and valid assessment of clinical skills, overcoming the limitations of traditional human grading. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=199 SRC="FIGDIR/small/26343786v1_ufig1.gif" ALT="Figure 1"> View larger version (58K): org.highwire.dtl.DTLVardef@e3990borg.highwire.dtl.DTLVardef@54c9a3org.highwire.dtl.DTLVardef@820d53org.highwire.dtl.DTLVardef@390508_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.