ChatGPT versus human in generating medical graduate exam questions - An international prospective study
Cheung, B. H. H.; Lau, G. K. K.; Wong, G. T. C.; Lee, E. Y. P.; Kulkarni, D.; Seow, C. S.; Wong, R.; Co, M. T. H.
Show abstract
IntroductionThis is a prospective study on the quality of multiple-choice questions (MCQs) generated by the language model ChatGPT for the use in medical graduate examination. Methods50 MCQs were generated by ChatGPT with reference to two standard undergraduate medical textbooks (Harrisons, and Bailey & Loves). Another 50 MCQs were drafted by two university professoriate staffs using the same medical textbooks. All 100 MCQ were individually numbered, randomized and sent to five independent international assessors for MCQ quality assessment using a standardized assessment score on five assessment domains; namely, appropriateness of the question, clarity and specificity, relevance, discriminative power of alternatives, and suitability for medical graduate examination. ResultsThe total time required for ChatGPT to create the 50 questions was 20 minutes 25 seconds while it took two human examiners a total of 211 minutes 33 seconds for drafting the 50 questions. When a comparison of the mean score was made between the questions constructed by AI with those drafted by human, only in the relevance domain that the AI was inferior to human (AI: 7.56 +/- 0.94 vs human: 7.88 +/- 0.52; p = 0.04). There was no significant difference in question quality between questions drafted by AI versus human, in the total assessment score as well as in other domains. Questions generated by AI yielded a wider range of scores while those created by human were consistent and within a narrower range. ConclusionChatGPT has the potential to generate comparable-quality MCQs for medical graduate examination within a significantly shorter time.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Large language models for generating medical examinations: systematic review 97%
- Performance of ChatGPT on Chinese National Medical Licensing Examinations: A Five-Year Examination Evaluation Study for Physicians, Pharmacists and Nurses 95%
- Medical students' perceptions towards artificial intelligence in education and practice: A multinational, multicenter cross-sectional study 94%
Similar papers in this journal
- Introducing the 4Ps Model of Transitioning to Distance Learning: a convergent mixed methods study conducted during the COVID-19 pandemic 94%
- Evaluation of Large Language Models in Medical Examinations:A Scoping Review Protocol 93%
- A national professional development program fills mentoring gaps for postdoctoral researchers 92%
Similar papers in this journal
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 94%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 93%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 93%
Similar papers in this journal
- Perceptions of Complementary, Alternative, and Integrative Medicine: Insights from a Large-Scale International Cross-Sectional Survey of Surgery Researchers and Clinicians 92%
- Cloud-controlled microscopy enables remote project-based biology education in Latinx communities in the United States and Latin America 91%
- Self-learning on COVID-19 among medical students and their preparedness to participate in government's COVID-19 response in Bhutan: a cross-sectional study 90%
Similar papers in this journal
- Evaluation of Self-Directed Learning Activities at King Abdulaziz University: A Qualitative Study of Faculty Perceptions 94%
- Validation of the patient reported outcome measures tool “Catquest” in Odia language 90%
- YouTube as an information source during the Coronavirus disease (COVID-19) pandemic: Evaluation of the Turkish and English content 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.