Closing the Paediatric Gap: Adult-Trained AI Generalises Robustly to Paediatric Coeliac Disease Diagnosis
Jaeckle, F.; Gillett, P. M.; Kirkwood, K. J.; Natu, S.; Chan, J. Y. H.; Bateman, A. C.; Arends, M. J.; Soilleux, E. J.
Show abstract
Background Coeliac disease (CD) diagnosis on duodenal biopsies is limited by interobserver variability. We have previously demonstrated pathologist-level performance with our artificial intelligence (AI) model for the histopathological diagnosis of adult CD, but not in paediatric practice. As paediatric CD screening programmes expand internationally, accurate and scalable diagnostic tools are needed. We investigated whether an AI model trained exclusively on adult whole-slide images (WSIs) can generalise to paediatric CD diagnosis across independent centres. Methods A training and validation dataset of 9,958 WSIs from 8,421 adult patients (961 CD) from five centres was used to develop an ensemble of multiple-instance learning models using features from a foundation model. Testing was performed on 708 consecutive paediatric patients (86 CD) from two centres (Edinburgh and Southampton) not included in training. Model calibration was assessed, and probability outputs were grouped into clinically interpretable categories. Findings In adult cross-validation, the AI model achieved an area under the receiver operating characteristic curve (AUC) of 98.7%, sensitivity of 84.9%, specificity of 99.0%, and negative predictive value (NPV) of 98.1%. On testing (paediatric) datasets, performance remained high (AUC 98.8%, sensitivity 80.2%, specificity 98.4%, NPV 97.3%). Restricting analysis to predictions outside the intermediate-probability range (predicted CD probability <10% or [≥]65%; 85.3% of cases) improved sensitivity to 100% and specificity to 98.7%. No misclassifications were observed among high-confidence predictions (<2% or [≥]85%; 66.0% of cases). The expected calibration error was 0.03. Performance improved significantly when biopsies from both duodenal sites (bulb [D1] and descending [D2/3]) were considered. Interpretation Our AI model, trained on adult biopsies, generalises to paediatric CD diagnosis across centres and scanner platforms. Well-calibrated probability outputs provide clinically interpretable measures of diagnostic confidence and could support safe identification of CD-negative biopsies within defined thresholds. These findings demonstrate the feasibility of applying adult-derived AI models in paediatric populations and reinforce the importance of multi-site (D1 & D2) biopsy sampling.
Matching journals
The top 11 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deep learning-based approach for the characterization and quantification of histopathology in mouse models of colitis 90%
- Abdominal Imaging Associates Body Composition with COVID-19 Severity 90%
- Colonic mucosal associated invariant T cells in Crohn’s disease have a diverse and non-public T cell receptor beta chain repertoire 89%
Similar papers in this journal
- Mucosal-Associated Invariant T (MAIT) Cells are Highly Activated in Duodenal Tissue of Humans with Vibrio cholerae O1 Infection 89%
- Clinical predictors for etiology of acute diarrhea in children in resource-limited settings 88%
- Environmental enteric dysfunction and small intestinal histomorphology of stunted children in Bangladesh 88%
Similar papers in this journal
- Machine Learning-Based Prediction of Pediatric Ulcerative Colitis Treatment Response using Diagnostic Histopathology 90%
- Clonal transitions and phenotypic evolution in Barrett esophagus 90%
- Crohn's patients and healthy infants share immunodominant B cell response to commensal flagellin peptide epitopes 88%
Similar papers in this journal
- Attention-based whole-slide image compression achieves pathologist-level pre-screening of multi-organ routine histopathology biopsies 90%
- MIXTURE of human expertise and deep learning—Developing an explainable model for predicting pathological diagnosis and survival in patients with interstitial lung disease 87%
- Tissue contamination challenges the credibility of machine learning models in real world digital pathology 86%
Similar papers in this journal
- AI portal tract detection and characterisation for a regional analysis of steatosis and inflammation in MASLD, MASH, and AIH 88%
- Extended laboratory panel testing in the Emergency Department for risk-stratification of patients with COVID-19: a single centre retrospective service evaluation 85%
- Lymphoid Enhancer-Binding Factor 1 (LEF1) immunostaining as a surrogate of β-catenin ( CTNNB1) mutations 84%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.