Are Llms Ready For Pediatrics? A Comparative Evaluation Of Model Accuracy Across Clinical Domains
Mondillo, G.; Colosimo, S.; Perrotta, A.; Frattolillo, V.; Masino, M.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWLarge Language Models (LLMs) are rapidly emerging as promising tools in the healthcare field, yet their effectiveness in pediatric contexts remains underexplored. This study evaluated the performance of eight contemporary LLMs, released between 2024 and 2025, in answering multiple-choice questions from the MedQA dataset, stratified into two distinct categories: adult medicine (1461 questions) and pediatrics (1653 questions). Models were tested using a standardized prompting methodology with default hyperparameters, simulating real-world use by non-expert clinical users. Accuracy scores for adult and pediatric subsets were statistically compared using the chi-square test with Yates correction. Five models (Amazon Nova Pro 1.0, GPT 3.5-turbo-0125, Gemini 2.0 Flash, Grok 2, and Claude 3 Sonnet) demonstrated significantly lower performance on pediatric questions, with accuracy drops of up to more than 10 percentage points. In contrast, ChatGPT-4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet showed comparable performance across both domains, with ChatGPT-4o achieving the most balanced result (accuracy: 83.57% adult, 83.18% pediatric; p = 0.80). These findings suggest that while some models struggle with pediatric-specific content, more recent and advanced LLMs may offer improved generalizability and domain robustness. The observed variability highlights the critical importance of domain-specific validation prior to clinical implementation, particularly in specialized fields such as pediatrics.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 95%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 95%
- Raising awareness of potential biases in medical machine learning: Experience from a Datathon 94%
Similar papers in this journal
- Knowledge Graph-based Thought: a knowledge graph enhanced LLMs framework for pan-cancer question answering 92%
- New implementation of data standards for AI research in precision oncology. Experience from EuCanImage 91%
- Strategies and Techniques for Quality Control and Semantic Enrichment with Multimodal Data: A Case Study in Colorectal Cancer with eHDPrep 91%
Similar papers in this journal
- Clinical code sets and the problem of redundancy in code set repositories 92%
- The Mastery Rubric for Bioinformatics: supporting design and evaluation of career-spanning education and training 92%
- Comparison of local large language models for extraction of signs and symptoms data from electronic health records 92%
Similar papers in this journal
- One LLM is not Enough: Harnessing the Power of Ensemble Learning for Medical Question Answering 96%
- Design and implementation of a system for automated monitoring of adherence to evidenced-based clinical guideline recommendations 93%
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 93%
Similar papers in this journal
- Automatic Detection and Extraction of Key Resources from Tables in Biomedical Papers 91%
- Expanding a Database-derived Biomedical Knowledge Graph via Multi-relation Extraction from Biomedical Abstracts 91%
- A compact encoding of the genome suitable for machine learning prediction of traits and genetic risk scores. 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.