Large Language Models in Healthcare Simulation Education: A Bibliometric Analysis with AI-Assisted Screening
Pears, M.; Wadhwa, K.; Payne, S. R.; Konstantinidis, S. T. H.; Biyani, C. S.
Show abstract
Large language models (LLMs) such as ChatGPT are rapidly reshaping healthcare education and simulation-based training in non-technical skills (NTS), yet no bibliometric analysis has mapped this landscape. We searched seven open-access databases (OpenAlex, PubMed, Europe PMC, Crossref, Semantic Scholar, CORE, DOAJ) for English-language publications from January 2020 to March 2026. From 100,277 initial records, a sequential keyword funnel yielded 830 candidate papers, which were screened by 83 independent Claude Sonnet 4.6 AI agents applying pre-specified inclusion criteria (PRISMA-trAIce compliant; Cohen's kappa = 0.86 pre-reconciliation, 1.0 post-reconciliation). The final AI-verified corpus comprised 551 papers with a compound annual growth rate of 109%, contributions from 2,398 authors across 279 journals in 58 countries, and an h-index of 41. ChatGPT dominated the model landscape (46% of papers), with open-source models virtually absent. Virtual patient chatbots were the leading simulation modality (106 papers). Among NTS domains, communication (145 papers) and decision-making (135 papers) were most studied, whereas teamwork, leadership, situational awareness, and crisis resource management were markedly underrepresented. Only 6 urology-relevant papers were identified, none examining LLM integration within boot camp training formats. The field is growing at extraordinary pace but remains concentrated in a narrow range of NTS domains and a single proprietary model. Critical gaps persist in team-based skills training, open-source model evaluation, and specialty-specific simulation. AI-assisted bibliometric screening using multiple independent agents is feasible, reliable, and scalable, offering a replicable methodology for mapping fast-evolving research fields.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Introducing the EMPIRE Index: A novel, value-based metric framework to measure the impact of medical publications 96%
- Investigating the Nature of Open Science Practices Across Complementary, Alternative, and Integrative Medicine Journals: An Audit 95%
- Transparency in peer review: Exploring the content and tone of reviewers' confidential comments to editors 94%
Similar papers in this journal
Similar papers in this journal
- Diversity and inclusion: A hidden additional benefit of Open Data 96%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 94%
- The NASSS (Non-Adoption, Abandonment, Scale-Up, Spread and Sustainability) framework use over time: A scoping review 93%
Similar papers in this journal
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 93%
- Impact of the Federated Data Platform's digital surgery scheduling system on elective theatre utilisation at an NHS Trust: an interrupted time series analysis 93%
- Network Graph Representation of COVID-19 Scientific Publications to Aid Knowledge Discovery 92%
Similar papers in this journal
- Emerging Applications of NLP and Large Language Models in Gastroenterology and Hepatology: A Systematic Review 91%
- Performance of o1 pro and GPT-4 in self-assessment questions for nephrology board renewal 91%
- Semantic and geographical analysis of Covid-19 trials reveals a fragmented clinical research landscape likely to impair informativeness. 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.