AI-Powered Test Question Generation in Medical Education: The DailyMed Approach
van Uhm, J.; van Haelst, M. M.; Jansen, P. R.
Show abstract
IntroductionLarge language models (LLMs) presents opportunities to improve the efficiency and quality of tools in medical education, such as the generation of multiple-choice questions (MCQs). However, ensuring that these questions are clinically relevant, accurate, and easily accesible and reusable remains challenging. Here, we developed DailyMed, an online automated pipeline using LLMs to generate high-quality medical MCQs. MethodsOur DailyMed pipeline involves several key steps: 1) topic generation, 2) question creation, 3) validation using Semantic Scholar, 4) difficulty grading, 5) iterative improvement of simpler questions, and 6) final human review. The Chain-of-Thought (CoT) prompting technique was applied to enhance LLM reasoning. Three state-of the art LLMs--OpenBioLLM-70B, GPT-4o, and Claude 3.5 Sonnet--were evaluated within the area of clinical genetics, and the generated questions were rated by clinical experts for validity, clarity, originality, relevance, and difficulty. ResultsGPT-4o produced the highest-rated questions, excelling in validity, originality, clarity, and relevance. Although OpenBioLLM was more cost-efficient, it consistently scored lower in all categories. GPT-4o also achieved the greatest topic diversity (89.8%), followed by Claude Sonnet (86.9%) and OpenBioLLM (80.0%). In terms of cost and performance, GPT-4o was the most efficient model, with an average cost of $0.51 per quiz and a runtime of 16 seconds per question. ConclusionsOur pipeline provides a scalable, effective and online-accessible solution for generating diverse, clinically relevant MCQs. GPT-4o demonstrated the highest overall performance, making it the preferred model for this task, while OpenBioLLM offers a cost-effective alternative.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Large Language Models in Real-World Clinical Workflows: A Systematic Review of Applications and Implementation 92%
- AI chatbots not yet ready for clinical use 92%
- Listening to mental health crisis needs at scale: using Natural Language Processing to understand and evaluate a mental health crisis text messaging service 91%
Similar papers in this journal
- Performance of ChatGPT on Chinese National Medical Licensing Examinations: A Five-Year Examination Evaluation Study for Physicians, Pharmacists and Nurses 93%
- Large language models for generating medical examinations: systematic review 92%
- Evaluation of Statistical Illiteracy in Latin American Clinicians and of the Efficacy of a 10-Hour Course 90%
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 93%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 93%
Similar papers in this journal
- Evaluation of the performance of GPT-3.5 and GPT-4 on the Medical Final Examination 94%
- CONSORT-TM: Text classification models for assessing the completeness of randomized controlled trial publications 93%
- The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.