Interpretable machine learning for coeliac disease diagnosis: quantitative morphometry of duodenal biopsies
Bryant, R.; Romero Diaz, J.; Scott, A. G.; Sagdeo, A. A.; Jenkins, G. Z.; Richardson, R. A.; Chan, J. Y. C.; Arends, M. J.; Soilleux, E. J.; Jaeckle, F.
Show abstract
Background Coeliac disease affects approximately 1% of the global population and remains substantially underdiagnosed. Histopathological assessment of duodenal biopsies is the diagnostic gold standard but is subject to approximately 20% inter-observer disagreement. While machine learning approaches show promise, most prior work relies on black-box models with limited interpretability, restricting clinical adoption. Methods We present an interpretable pipeline that follows established histopathological criteria by extracting clinically meaningful morphological features from H&E-stained whole-slide images. Five sequential stages perform pre-processing, semantic segmentation of villi, crypts, intraepithelial lymphocytes (IELs) and enterocytes, crypt morphometry, villus length estimation via a novel polyline-based keypoint model, and coeliac disease classification using three quantitative features: IEL-to-enterocyte ratio, villus-to-crypt area ratio, and villus-length-to-crypt-depth ratio. Training and validation used data from four institutions; independent testing used 1,357 WSIs from two further institutions including one with a previously unseen scanner manufacturer, spanning five diagnostic categories: coeliac disease, normal mucosa, chronic inflammation, gastric metaplasia, and gastric heterotopia. Results Semantic segmentation achieved villus and crypt precision and recall of 87-90%. Villus length estimation correlated strongly with expert annotations (Pearson's r=0.85, mean relative error 13.5% post-calibration). All three morphological features significantly separated coeliac disease from all non-coeliac diagnostic groups across internal and external datasets (p<0.01 in all comparisons). On the test set the diagnostic classifier achieved accuracy 94.5%, PPV 92.9%, NPV 94.7%, and AUC 0.982. Conclusions This interpretable framework achieves strong multi-centre diagnostic performance while producing quantitative morphological outputs, villus length, crypt depth, and IEL-to-enterocyte ratios, that directly reflect established histopathological criteria, representing a meaningful step towards standardised AI-assisted coeliac disease diagnosis.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ROSIE: AI generation of multiplex immunofluorescence staining from histopathology images 93%
- A SIMPLI (Single-cell Identification from MultiPLexed Images) approach for spatially resolved tissue phenotypingat single-cell resolution. 93%
- METI: Deep profiling of tumor ecosystems by integrating cell morphology and spatial transcriptomics 92%
Similar papers in this journal
- Attention-based whole-slide image compression achieves pathologist-level pre-screening of multi-organ routine histopathology biopsies 95%
- Interpretable multimodal deep learning for real-time pan-tissue pan-disease pathology search on social media 93%
- MIXTURE of human expertise and deep learning—Developing an explainable model for predicting pathological diagnosis and survival in patients with interstitial lung disease 91%
Similar papers in this journal
- Fostering transparent medical image AI via an image-text foundation model grounded in medical literature 92%
- Evaluating and Mitigating Limitations of Large Language Models in Clinical Decision Making 92%
- 3D spatially-resolved geometrical and functional models of human liver tissue reveal new aspects of NAFLD progression 91%
Similar papers in this journal
Similar papers in this journal
- Deep learning-based approach for the characterization and quantification of histopathology in mouse models of colitis 93%
- Deep Learning Classification of Lipid Droplets in Quantitative Phase Images 91%
- LIMPACAT : Multi-Omics Attention Transformer for Immune Prediction in Liver Cancer Using Whole-Slide Imaging 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.