Automatic Physical Examination Segmentation within Objective Structured Clinical Examination Videos
Kang, S.; Holcomb, M. J.; Hein, D.; Shakur, A. H.; Dalton, T. O.; Jamieson, A. R.
Show abstract
ObjectiveAssessing medical student performance in Objective Structured Clinical Examinations (OSCEs) is labor-intensive, requiring trained evaluators to review 15-minute long videos. The physical examination period constitutes only a small portion of these videos. Automated segmentation of OSCE videos could significantly streamline the evaluation process by detecting this physical exam portion for targeted evaluation. Current video analysis approaches struggle with these long recordings due to computational constraints and challenges in maintaining temporal context. This study tests whether multimodal large language models (MM-LLMs) can segment physical examination periods in OSCE videos without prior training, potentially easing the burden on both human graders and automated systems. MethodsWe analyzed 500 videos from five OSCE stations at UT Southwestern Simulation Center, each 15 minutes long, using hand-labeled physical examination periods as ground truth. MM-LLMs processed video frames at one frame per second, classifying them into discrete activity states. A hidden Markov model with Viterbi decoding ensured temporal consistency across segments, addressing the inherent challenges of frame-by-frame classification. ResultsUsing Viterbi decoding trained on just 50 hand-labeled videos (10 from each station), zero-shot GPT-4o achieved 99.8% recall and 78.3% intersection over union (IOU), effectively capturing physical examinations with an average duration of 175 seconds from 900-second videos--an 81% reduction in frames requiring review. ConclusionsIntegrating multimodal large language models with temporal modeling effectively segments physical examination periods in OSCE videos without requiring extensive training data. This approach significantly reduces review time while maintaining clinical assessment integrity, demonstrating that zero-shot AI methods can be optimized for medical educations specific requirements. The technique establishes a foundation for more efficient and scalable clinical skills assessment across diverse medical education settings.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 94%
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 94%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 93%
Similar papers in this journal
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 93%
- A Deep Learning Based Smartphone Application for Early Detection of Nasopharyngeal Carcinoma Using Endoscopic Images 92%
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 91%
Similar papers in this journal
- One LLM is not Enough: Harnessing the Power of Ensemble Learning for Medical Question Answering 92%
- Design and implementation of a system for automated monitoring of adherence to evidenced-based clinical guideline recommendations 92%
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 91%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.