Pneumonia Detection in Paediatric Chest X-Rays using Ensembled Large Language Models
Tan, J.; Tang, P. H.
Show abstract
BackgroundPa4ediatric pneumonia is a major cause of childhood morbidity and mortality. Chest X-rays (CXR) are central to diagnosis, but shortages of specialist radiologists can delay reporting. Multimodal large language models (MLLMs) may assist clinical workflows by analysing images and communicating findings, although their diagnostic performance remains below state-of-the-art classifiers. ObjectiveTo evaluate whether ensemble strategies improve MLLM diagnostic performance for paediatric radiological pneumonia detection on CXRs. MethodsIn this retrospective study, paediatric CXRs from two datasets (balanced and real-world) at KK Womens and Childrens Hospital were analysed. Images were independently reviewed by two board-certified radiologists, with pneumonia severity assigned to three classes using a predefined consensus algorithm. Fifteen MedGemma-4B-it agents classified each CXR into five likelihood categories, which were mapped to the three severity classes for evaluation. Majority voting, soft voting and GPTOSS-20B aggregation were compared with baseline average agent performance. The primary outcome was One-vs-Rest (OvR) AUROC. Secondary metrics included accuracy, sensitivity, specificity, F1-score, Cohens {kappa} and One-vs-One (OvO) AUROC. ResultsThe balanced dataset contained 900 CXRs and the real-world dataset 1300 CXRs. Soft voting significantly improved OvR-AUROC compared with baseline in both datasets (Balanced: 0.829>0.764; 95%CI=0.752-0.779; P=0.0002. Real-world: 0.728>0.655; 95%CI=0.638-0.679; P=0.0003). Soft voting also improved accuracy, Cohens {kappa}, OvO-AUROC in both datasets and F1-score in the balanced dataset. ConclusionSoft voting enhances MedGemmas diagnostic discriminatory performance for paediatric radiological pneumonia detection. Our system enables privacy-preserving, near real-time clinical decision support with explainable outputs, having potential for integration into emergency departments. Our systems high specificity supports triage by flagging high-risk radiological pneumonia cases. Clinical ImpactO_LIPaediatric CXRs often face reporting delays exceeding 24 hours due to radiologist shortages. C_LIO_LIOur proposed MLLM ensemble framework achieves better than average MLLM diagnostic discrimination for radiological pneumonia without requiring cloud-based systems. C_LIO_LISoft-voting aggregation enhances diagnostic discriminatory effectiveness for paediatric pneumonia severity, while preserving explainable outputs. C_LIO_LIOur system acts as a decision support tool that identifies higher-risk pneumonia cases for urgent review, supporting safer triage. C_LI
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and validation of AI-based pre-screening of large bowel biopsies 93%
- Multicenter Validation of a Machine Learning Algorithm for Diagnosing Pediatric Patients with Multisystem Inflammatory Syndrome and Kawasaki Disease 92%
- Novel deep learning algorithm predicts the status of molecular pathways and key mutations in colorectal cancer from routine histology images 91%
Similar papers in this journal
- An ML prediction model based on clinical parameters and automated CT scan features for COVID-19 patients 95%
- Toward Understanding COVID-19 Pneumonia: A Deep-learning-based Approach for Severity Analysis and Monitoring the Disease 94%
- High-Dimensional Multinomial Multiclass Severity Scoring of COVID-19 Pneumonia Using CT Radiomics Features and Machine Learning Algorithms 94%
Similar papers in this journal
- Development and Validation of a Deep Learning Model for Detecting Signs of Tuberculosis on Chest Radiographs among US-bound Immigrants and Refugees 95%
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 94%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 94%
Similar papers in this journal
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 93%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 93%
- The clinician-AI interface: intended use and explainability in FDA-cleared AI devices for medical image interpretation 93%
Similar papers in this journal
- Protocol for the development and validation of a machine-learning based tool for predicting the risk of hypertriglyceridemia in critically-ill patients receiving propofol sedation 92%
- Development and validation of multivariable machine learning algorithms to predict risk of cancer in symptomatic patients referred urgently from primary care 90%
- Developing a video expert panel as a reference standard to evaluate respiratory rate counting in paediatric pneumonia diagnosis: protocol for a cross-sectional study 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.