Characterizing large language model generative artificial intelligence variability in the production of objective structured clinical examination stations
Joseph-Delaffon, K.; Desgrouas, M.; Catanese, S.; Lejeune, J.; Nait-Kaci, J.; Piver, E.; Breteau, I.; Leducq, S.; Gatault, P.; Khanna, R. K.; Angoulvant, D.; Vallet, N.
Show abstract
Background. Designing high-quality Objective Structured Clinical Examination (OSCE) stations is a time-consuming process. Generative artificial intelligence (AI) represents a promising path to accelerate content creation by automating the generation of scenarios. A growing number of AI tools is now available for this purpose. Objective. To assess the variability between generative AI models in their ability to produce OSCE stations in the field of paediatrics. Methods. A structured prompt was developed based on the French national OSCE guidelines for medical education. Five distinct AI models were provided with this prompt, alongside the neonatal jaundice chapter from the French pediatric reference textbook, to generate 6 complete OSCE stations. Results. Prompt compliance was high for ChatGPT 5.1, ChatGPT 5.2, Gemini 3.0 Pro, and Claude Opus 4.5, while it was lower for Grok 4.1. Expert-rated quality was generally high, with few factual errors or missing information across models. However usability differed significantly between models. This was also true for several quality dimensions such as checklist clarity, embedding of checklist answers within vignettes, and ease of standardized patient formation. ChatGPT 5.1 required the most revisions and Gemini most often rated usable as is. Significant inter-model differences were observed in diagnostics, only with ChatGPT 5.1 sampling all three neonatal jaundice categories. Contextual variables showed systematic narrowing across models. Clinical grid density was consistent (10-12 items per station), but thematic distribution differed markedly. Soft skills coverage varied significantly across models (p=0.002), none of them consistently representing all communication competency domains. Conclusion. Large language models can generate structurally compliant OSCE stations, but surface compliance conceals substantive inter-model differences in diagnostic coverage, contextual diversity, and soft skills representation, that compromise content validity. No model currently meets the criteria for unsupervised deployment in a summative assessment bank. The choice of model carries pedagogical implications and expert curation remains essential before integration into high-stakes assessment workflows.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Large language models for generating medical examinations: systematic review 93%
- Performance of ChatGPT on Chinese National Medical Licensing Examinations: A Five-Year Examination Evaluation Study for Physicians, Pharmacists and Nurses 92%
- Medical students' perceptions towards artificial intelligence in education and practice: A multinational, multicenter cross-sectional study 91%
Similar papers in this journal
- Medical Clinical Minds Meet Artificial Intelligence: Italian Physicians' Knowledge, Attitudes, and Concordance between Italian Physicians and AI-Generated Diagnoses. A National Cross-Sectional Study 92%
- Human-supervised, large language model-based clinical decision support aligned to national newborn protocols in Kenya: a pragmatic, early-stage evaluation 91%
- Large Language Models in Real-World Clinical Workflows: A Systematic Review of Applications and Implementation 90%
Similar papers in this journal
- Evaluation of Large Language Models in Medical Examinations:A Scoping Review Protocol 93%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 92%
- Evaluating Large Language Models for ADHD Education: A Comparative Study of ChatGPT 5, DeepSeek V3, and Grok 4 91%
Similar papers in this journal
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 94%
- Machine Learning for Paediatric Related Decision Support in Emergency Care - A UK and Ireland Network Survey Study 93%
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 93%
Similar papers in this journal
- Evaluation of the performance of GPT-3.5 and GPT-4 on the Medical Final Examination 94%
- Development and testing of a game-based digital intervention for working memory training in autism spectrum disorder 90%
- Large Language Models for Zero-Shot Procedure Extraction in Orthopedic Surgery: A Comparative Evaluation 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.