Quality of Human Expert vs. Large Language Model Generated Multiple Choice Questions in the Field of Mechanical Ventilation
Safadi, S.; Amirahmadi, R.; Tlimat, A.; Rovinski, R.; Sun, J.; Lee, B. W.; Seam, N.
Show abstract
BackgroundMechanical ventilation (MV) is a critical competency in critical care training, yet standardized methods for assessing MV-related knowledge are lacking. Traditional multiple-choice question (MCQ) development is resource-intensive, and prior studies have suggested that generative AI tools could streamline question creation. However, the effectiveness and reliability of AI- generated MCQs remain unclear. This study evaluates whether MCQs generated by ChatGPT are non-inferior to human-expert (HE) created questions in terms of quality and relevance for MV education. MethodsThree key MV topics were selected: Equation of Motion & Ohms Law, Tau & Auto PEEP, and Oxygenation. Fifteen learning objectives were used to generate 15 AI-written MCQs via a standardized prompt with ChatGPT o1 (model o1-preview-2024-09-12). A group of 31 faculty experts, all of whom regularly teach MV, evaluated both AI-generated and HE-generated MCQs. Each MCQ was assessed based on its alignment with learning objectives, accuracy, clarity, plausibility of distractors, and difficulty level. The faculty members were blinded to the provenance of the MCQ questions. The non-inferiority margin was predefined as 15% of the total possible score (-3.45). ResultsAI-generated MCQs were statistically non-inferior to expert-written MCQs (95% upper CI: [-1.15, {infty}]). Additionally, respondents were unable to reliably differentiate AI-generated from HE-written MCQs (p = 0.32). ConclusionAI-generated MCQs using ChatGPT o1 are comparable in quality and difficulty to those written by human experts. Given the time and resource-intensive nature of human MCQ development, AI-assisted question generation may serve as an efficient and scalable alternative for medical education assessment, even in highly specialized domains such as mechanical ventilation.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Large language models for generating medical examinations: systematic review 97%
- Medical students' perceptions towards artificial intelligence in education and practice: A multinational, multicenter cross-sectional study 94%
- Student self-assessment: feasibility, advantages and limitations Example of a workshop for trainee surgeons using a suture score 94%
Similar papers in this journal
- Evaluating user experience with immersive technology in simulation-based education: a modified Delphi study with qualitative analysis 95%
- Introducing the 4Ps Model of Transitioning to Distance Learning: a convergent mixed methods study conducted during the COVID-19 pandemic 94%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 94%
Similar papers in this journal
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 94%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 93%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 93%
Similar papers in this journal
- Handling and Packaging of Medical Bags at the Acute Disaster Site Under High Temperature Conditions 88%
- Impact of COVID-19 Upon Changes in Emergency Room Visits with Chest Pain of Possible Cardiac Origin 87%
- Occupational safety and health aspects of corporate social responsibility reporting in Japan: comparison between 2012 and 2020 87%
Similar papers in this journal
- Performance of Advanced Large Language Models (GPT-4o, GPT-4, Gemini 1.5 Pro, Claude 3 Opus) on Japanese Medical Licensing Examination: A Comparative Study 95%
- Synthetic Data Generation in Healthcare: A Scoping Review of reviews on domains, motivations, and future applications 92%
- Image and structured data analysis for prognostication of health outcomes in patients presenting to the Emergency Department during the COVID-19 pandemic 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.