AI Video Analysis of Psychomotor Performance in EMS Education: Agreement With Human Evaluators Across Three Skills
Otte, J. H.; Cartagena, A.
Show abstract
Background. A primary constraint on the capacity of EMS programs to meet industry demand is psychomotor instruction and verification, requiring direct observation of each student by a qualified evaluator. Whether AI video analysis can relieve it is untested; none has been applied to EMS skill examination or compared with human examiners. Objective. To quantify human EMS evaluator inter-rater reliability and evaluate an AI video-analysis platform against it. Methods. In a prospective, fully crossed study, five certified EMS evaluators and an AI platform independently scored identical video-recorded EMT performances of cervical collar application (n=15), bag-valve-mask (BVM) ventilation (n=14), and medical assessment (n=15) on dichotomous checklists with critical-failure criteria. Agreement was assessed at item, score, and decision levels using Fleiss' kappa, Krippendorff's alpha, Gwet's AC1, and ICC(2,1)/ICC(2,k). Results. Human item agreement was moderate (kappa 0.409 to 0.467), as was single-rater reliability (ICC(2,1) 0.539 to 0.694), against good panel reliability (ICC(2,k) 0.854 to 0.919). Recorded pass/fail agreement was fair (kappa 0.297 to 0.388) and critical-failure agreement near zero for two skills (kappa 0.028, 0.119). AI alignment tracked rubric observability rather than task complexity: r = 0.857 (collar, exceeding every human), -0.173 (BVM), 0.664 (medical), and it was most lenient on two skills. Conclusions. Human evaluators are an imperfect standard, especially on critical failures. The AI was a legitimate additional rater where checklist items were discrete and visually verifiable, but not where credit required judging continuous quantities such as ventilation rate, volume, or suction duration. Defensible uses are formative and archival, not summative. These results reflect an early, non-specialist configuration: a baseline, not a limit.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating N95 Respirator Designs: A Mixed-Methods Pilot and Feasibility Study 92%
- Identifying clinical skill gaps of healthcare workers using a digital clinical decision support algorithm during outpatient pediatric consultations in primary health centers in Rwanda 92%
- Evaluating user experience with immersive technology in simulation-based education: a modified Delphi study with qualitative analysis 91%
Similar papers in this journal
- Improving capacity for advanced training in obstetric surgery: Evaluation of a blended learning approach 92%
- Team-Based Learning Versus Lecture-Based Instruction for Chest Radiograph Interpretation in Physician Associate Education: A Quasi-Experimental Study 92%
- Student self-assessment: feasibility, advantages and limitations Example of a workshop for trainee surgeons using a suture score 92%
Similar papers in this journal
- Medication errors during simulated paediatric resuscitations: a prospective, observational human reliability analysis 93%
- COVID-19 Advanced Respiratory Care Educational Training Program for Healthcare Workers in Lesotho: An Observational Study 93%
- A mixed methods study protocol to develop and pilot a Competency Assessment Tool to support therapists in the care of patients with blunt CHest trauma (CATCh Study) 92%
Similar papers in this journal
- Improving communications in PPE: A solution for ‘landline’ telephone communication 91%
- Diversity of CPR manikins for basic life support education: Use of manikin sex, race, and body shape – A scoping review 91%
- The Balancing Act of Academic Clinical Fellows in UK Emergency Medicine: A Qualitative Study 91%
Similar papers in this journal
- A Modified Delphi Consensus-based Comprehensive Checklist and Angoff Standard for Assessment of Competency in Brain Death/Death by Neurologic Criteria Determination 93%
- An international factorial vignette-based survey of intubation decisions in acute hypoxemic respiratory failure 91%
- Prone Positioning in a North American Cohort of Hypoxemic Patients on Mechanical Ventilation 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.