What level of expertise is necessary to generate ACLS training test questions: pre-med students vs. artificial intelligence?
LoGalbo, S. S.; Richman, M.; Wang, J.; Saji, I.; Traore, A.; Oliva, H.; Wu, E.; Drudi, A.; Foster, D.; Bhandari, S.; Delfillo, R. L.; McCann, A.; Coard, J.; Matthew, C.; Smith, B.
Show abstract
Abstract Introduction In-hospital cardiac arrest carries high mortality despite standardized ACLS training. Educators face increasing time constraints in developing assessment tools for ACLS training. Two possible solutions to this problem are using pre-medical students or using artificial intelligence to generate test questions. This study compared the quality of pre-medical student-generated ACLS test questions vs. AI-generated ACLS test questions, testing the hypothesis that AI-generated questions are non-inferior to student-generated questions. Methods Ten pre-medical students created ACLS questions following predefined criteria, while an AI model (Northwell's Artificial Intelligence Hub) generated comparable questions. A blinded ACLS-certified physician evaluated questions on the qualities of Alignment, Clarity, Cognitive Level, and Question Design using a standardized rubric (Likert scale: 1 = poor quality, 5 = excellent). Student's T-test and Chi-square analysis were used to compare the quality of questions on different rubric domains within each arm (student vs. AI) and within one domain (eg, question Clarity) between arms. The Student's T test was used when 2 comparator groups were compared (eg, Clarity of student-generated vs. AI-generated questions) within one arm. The ANOVA test was used when comparing more than 2 comparator groups (eg, Alignment vs. Clarity vs. Cognitive Level) within one arm. Statistical significance was set as a priority at p <0.05. Results Both student-generated and AI-generated questions were of high quality. AI-generated questions achieved the maximum score in the domains of Alignment, Clarity, and Question Design, but fell short of perfect scores in the domain of Cognitive Level (8 of 50 questions were less than 5). Student-generated questions achieved less-than-perfect scores in each domain. No significant difference was found in overall mean question scores between groups (students = 4.79, AI = 4.81; p = 0.9). However, AI-generated questions had significantly-greater Clarity (students = 4.8, AI = 5; p = .0461), while Alignment, Cognitive level, and Question Design showed no significant differences. Conclusion AI-generated questions demonstrated overall quality comparable to those generated by pre-medical students, supporting the potential role of AI as a scalable tool in ACLS educational assessment development. Further studies are warranted to evaluate additional AI platforms and determine optimal integration of AI in medical education assessment design.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Large language models for generating medical examinations: systematic review 96%
- Evaluation of Statistical Illiteracy in Latin American Clinicians and of the Efficacy of a 10-Hour Course 95%
- Student self-assessment: feasibility, advantages and limitations Example of a workshop for trainee surgeons using a suture score 94%
Similar papers in this journal
- Evaluating user experience with immersive technology in simulation-based education: a modified Delphi study with qualitative analysis 94%
- Introducing the 4Ps Model of Transitioning to Distance Learning: a convergent mixed methods study conducted during the COVID-19 pandemic 94%
- Clinical code sets and the problem of redundancy in code set repositories 94%
Similar papers in this journal
- Evaluation of Self-Directed Learning Activities at King Abdulaziz University: A Qualitative Study of Faculty Perceptions 93%
- Hospitalist perspectives of available tests to monitor volume status in patients with heart failure: a qualitative study 92%
- Validation of the patient reported outcome measures tool “Catquest” in Odia language 90%
Similar papers in this journal
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 94%
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 94%
- Harnessing the Open Access Version of ChatGPT for Enhanced Clinical Opinions 94%
Similar papers in this journal
- Performance of digital Early Warning Score (NEWS2) in a cardiac specialist setting: retrospective cohort study 93%
- Evaluating a first fully automated interview grounded in Multiple Mini Interview (MMI) methodology: results from a feasibility study 93%
- A mixed methods study protocol to develop and pilot a Competency Assessment Tool to support therapists in the care of patients with blunt CHest trauma (CATCh Study) 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.