Cancer-Tissue Fraction as a Scanner-Robust Triage Signal for Automated Gleason Grading of Prostate Biopsies: External Validation Across a Middle Eastern Cohort
Ebbert, J. L.; Perry, A.; Szymanski, J.; Della Corte, D.
Show abstract
Background: Deep-learning systems for Gleason grading are developed almost entirely on high-end clinical scanners and on cohorts from a small number of Western institutions, yet deployment increasingly involves other devices and other populations. These two distribution shifts, device and population, are rarely tested together on the same physical slides. The PAR dataset, from Erbil, Iraq, digitizes each biopsy on three scanners and provides three distinct pathologist grades, so it permits both tests at once on a Middle Eastern cohort. A concurrent study by the dataset originators validated a task-specific model and two foundation models on PAR; we complement it by testing an inde-pendently developed detect-then-grade pipeline and by separating scanner effects on detection from scanner effects on grading. Methods: We applied one fixed model de-veloped on North American and European material to all 1017 whole-slide images (339 slides from 185 patients, three scanners; 49.6% clinically significant cancer) with no scanner-specific or population-specific tuning. We measured cancer detection (area under the ROC curve of the predicted cancer-tissue fraction), all-slide ISUP agreement of the deployed detect-then-grade pipeline (quadratic-weighted kappa, QWK), and grading agreement on pathologist-confirmed cancers, at the slide level and, because a case carries up to two slides, at the patient level. The reference reader was S.A.; thresholds and operating points were cross-validated leave-one-out; scanners were compared by paired within-biopsy bootstrap and confidence intervals confirmed by patient-cluster bootstrap. Results: Detection was statistically equivalent across scanners (AUC 0.987 to 0.991; paired differences at most 0.003) and transferred to this non-Western cohort with no per-population tuning. At a 95% sensitivity operating point the deployed pipeline reached cross-validated all-slide QWK of 0.86, 0.81, and 0.86 (Grundium, Hamamatsu, Leica), matching the inter-pathologist ceiling of 0.81, against 0.23 to 0.62 for the ungated model. Grading of confirmed cancers was scanner dependent: the compact Grundium (0.63) did not differ from the clinical Leica (0.67; paired difference 0.04, 95% CI -0.03 to 0.11), while both exceeded Hamamatsu (0.44). Results held at the patient level, with grading somewhat lower for two scanners; the two slides of a case disagreed in grade in 43% of cases, and patient clustering did not widen the intervals. Conclusions: Can-cer-tissue fraction is a triage signal robust across scanner and transferable to an un-derrepresented population for detection, while grading is the scanner-sensitive step. Prostate grading models should be deployed as a detect-then-grade pipeline, with grading validated per device and confirmed on the local population.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Clinical-Grade Validation of an Autofluorescence Virtual Staining System with Human Experts and a Deep Learning System for Prostate Cancer 96%
- Artificial Intelligence for Advance Requesting of Immunohistochemistry in Diagnostically Uncertain Prostate Biopsies 96%
- Tissue contamination challenges the credibility of machine learning models in real world digital pathology 94%
Similar papers in this journal
- Automated Clear Cell Renal Carcinoma Grade Classification with Prognostic Significance 92%
- Classification performance bias between training and test sets in a limited mammography dataset 91%
- FalseColor-Python: a rapid intensity-leveling and digital-staining package for fluorescence-based slide-free digital pathology 91%
Similar papers in this journal
- Development and validation of AI-based pre-screening of large bowel biopsies 94%
- Novel deep learning algorithm predicts the status of molecular pathways and key mutations in colorectal cancer from routine histology images 90%
- Multicenter Validation of a Machine Learning Algorithm for Diagnosing Pediatric Patients with Multisystem Inflammatory Syndrome and Kawasaki Disease 87%
Similar papers in this journal
- Public evidence on AI products for digital pathology 92%
- A human-in-the-loop explanation framework for morphologically transparent AI predictions from whole-slide images 91%
- Diagnostic Accuracy of Artificial Intelligence in Classifying HER2 Status in Breast Cancer Immunohistochemistry Slides and Implications for HER2-Low Cases: A Systematic Review and Meta-Analysis 91%
Similar papers in this journal
- Independent assessment of a deep learning system for lymph node metastasis detection on the Augmented Reality Microscope 93%
- Bladder Cancer Prognosis Using Deep Neural Networks and Histopathology Images 92%
- Deep learning accurately quantifies plasma cell percentages on CD138-stained bone marrow samples 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.