Understanding Human AI Discrepancy in Breast Cancer TIL Assessment: A Multi-Rater and Perceptual Bias Study
Capar, A.; Aloglu, I.; Aker, F.; Ertano, M.; Mese, Y. E.; Ungor, A.; Yildiz, B. E.
Show abstract
Objective: Tumor-infiltrating lymphocytes (TILs) in breast cancer are one of the most important indicators of the immune response within the tumor microenvironment. They play a particularly significant prognostic and predictive role in triple-negative and HER2-positive subtypes. However, substantial inter-observer variability has been reported in TIL scoring among pathologists, which limits its reliability in clinical practice. The aim of this study was to evaluate the agreement between artificial intelligence (AI) models and pathologists in TIL scoring and to compare this agreement using different statistical approaches, thereby assessing the potential of AI integration into pathology practice. Materials and Methods: Digitized histopathological images of breast cancer cases were included in the study. Tumor regions annotated by pathologists were evaluated for both stromal TIL percentage and the proportion of stromal tumor area within each ROI, with assessments performed independently by three pathologists and two AI models. Agreement was assessed among pathologists, between pathologists and AI, and between AI models. Statistical analyses included intraclass correlation coefficient (ICC), Cohen and Fleiss kappa, correlation tests, and Bland-Altman analysis. In addition, categorical agreement was examined using different cut-off values. Results: Inter-pathologist agreement was high, with an ICC of 0.81. In contrast, the global agreement between pathologists and AI models was lower (ICC 0.41). Pairwise comparisons of pathologist-AI agreement yielded substantially lower ICC values (0.12-0.21), although this improved to 0.53 when three pathologists were assessed jointly with a single AI model. The strongest categorical agreement was observed with dichotomized TIL scores ([≤]10% vs. >10%), whereas multi-category classifications were associated with a marked reduction in kappa values. Spearman correlation coefficients between pathologists and AI models ranged from moderate to good ({rho} = 0.48-0.81). Agreement between the two AI models themselves was moderate, with an ICC of 0.64
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Independent assessment of a deep learning system for lymph node metastasis detection on the Augmented Reality Microscope 96%
- Deep learning accurately quantifies plasma cell percentages on CD138-stained bone marrow samples 94%
- Creating virtual H&E images using samples imaged on a commercial CODEX platform 94%
Similar papers in this journal
- Clinical-Grade Validation of an Autofluorescence Virtual Staining System with Human Experts and a Deep Learning System for Prostate Cancer 95%
- Artificial Intelligence for Advance Requesting of Immunohistochemistry in Diagnostically Uncertain Prostate Biopsies 94%
- Attention-based whole-slide image compression achieves pathologist-level pre-screening of multi-organ routine histopathology biopsies 94%
Similar papers in this journal
- Development of a multi-scanner facility for data acquisition for digital pathology artificial intelligence 94%
- A digital score of peri-epithelial lymphocytic activity predicts malignant transformation in oral epithelial dysplasia 93%
- Spatial Effects of Infiltrating T cells on Neighbouring Cancer Cells and Prognosis in Stage III CRC patients 90%
Similar papers in this journal
- Weakly supervised learning for multi-organ adenocarcinoma classification in whole slide images 94%
- Pixelwise H-score: a novel digital image analysis based-metric to quantify membrane biomarker expression from immunohistochemistry images 94%
- Classification performance bias between training and test sets in a limited mammography dataset 93%
Similar papers in this journal
- Automated and Manual Quantification of Tumour Cellularity in Digital Slides for Tumour Burden Assessment 95%
- PathProfiler: Automated Quality Assessment of Retrospective Histopathology Whole-Slide Image Cohorts by Artificial Intelligence, A Case Study for Prostate Cancer Research 95%
- VISTA: Virtual ImmunoSTAining for pancreatic disease quantification in murine cohorts 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.