iQC: machine-learning-driven prediction of surgical procedure uncovers systematic confounds of cancer whole slide images in specific medical centers
Schaumberg, A. J.; Lewis, M. S.; Nazarian, R.; Wadhwa, A.; Kane, N.; Turner, G.; Karnam, P.; Devineni, P.; Wolfe, N.; Kintner, R.; Rettig, M. B.; Knudsen, B. S.; Garraway, I. P.; Pyarajan, S.
Show abstract
ProblemThe past decades have yielded an explosion of research using artificial intelligence for cancer detection and diagnosis in the field of computational pathology. Yet, an often unspoken assumption of this research is that a glass microscopy slide faithfully represents the underlying disease. Here we show systematic failure modes may dominate the slides digitized from a given medical center, such that neither the whole slide images nor the glass slides are suitable for rendering a diagnosis. MethodsWe quantitatively define high quality data as a set of whole slide images where the type of surgery the patient received may be accurately predicted by an automated system such as ours, called "iQC". We find iQC accurately distinguished biopsies from nonbiopsies, e.g. prostatectomies or transurethral resections (TURPs, a.k.a. prostate chips), only when the data qualitatively appeared to be high quality, e.g. vibrant histopathology stains and minimal artifacts. Crucially, prostate needle biopsies appear as thin strands of tissue, whereas prostatectomies and TURPs appear as larger rectangular blocks of tissue. Therefore, when the data are of high quality, iQC (i) accurately classifies pixels as tissue, (ii) accurately generates statistics that describe the distribution of tissue in a slide, and (iii)accurately predicts surgical procedure from said statistics. We additionally compare our "iQC" to "HistoQC", both in terms of how many slides are excluded and how much tissue is identified in the slides. ResultsWhile we do not control any medical centers protocols for making or storing slides, we developed the iQC tool to hold all medical centers and datasets to the same objective standard of quality. We validate this standard across five Veterans Affairs Medical Centers (VAMCs) and the Automated Gleason Grading Challenge (AGGC) 2022 public dataset. For our surgical procedure prediction task, we report an Area Under Receiver Operating Characteristic (AUROC) of 0.9966-1.000 at the VAMCs that consistently produce high quality data and AUROC of 0.9824 for the AGGC dataset. In contrast, we report an AUROC of 0.7115 at the VAMC that consistently produced poor quality data. An attending pathologist determined poor data quality was likely driven by faded histopathology stains and protocol differences among VAMCs. Corroborating this, iQCs novel stain strength statistic finds this institution has significantly weaker stains (p < 2.2 x 10-16, two-tailed Wilcoxon rank-sum test) than the VAMC that contributed the most slides, and this stain strength difference is a large effect (Cohens d = 1.208). In addition to accurately detecting the distribution of tissue in slides, we find iQC recommends only 2 of 3736 VAMC slides (0.005%) be reviewed for inadequate tissue. With its default configuration file, HistoQC excluded 89.9% of VAMC slides because tissue was not detected in these slides. With our customized configuration file for HistoQC, we reduced this to 16.7% of VAMC slides. Strikingly, the default configuration of HistoQC included 94.0% of the 1172 prostate cancer slides from The Cancer Genome Atlas (TCGA), which may suggest HistoQC defaults were calibrated against TCGA data but this calibration did not generalize well to non-TCGA datasets. For VAMC and TCGA, we find a negligible to small degree of agreement in the include/exclude status of slides, which may suggest iQC and HistoQC are not equivalent. ConclusionOur surgical procedure prediction AUROC may be a quantitative indicator positively associated with high data quality at a medical center or for a specific dataset. We find iQC accurately identifies tissue in slides and excludes few slides, unless the data are poor quality. To produce high quality data, we recommend producing slides using robotics or other forms of automation whenever possible. We recommend scanning slides digitally before the glass slide has time to develop signs of age, e.g faded stains and acrylamide bubbles. We recommend using high-quality reagents to stain and mount slides, which may slow aging. We recommend protecting stored slides from ultraviolet light, from humidity, and from changes in temperature. To our knowledge, iQC is the first automated system in computational pathology that validates data quality against objective evidence, e.g. surgical procedure data available in the EHR or LIMS, which requires zero efforts or annotations from anatomic pathologists. Please see https://github.com/schaumba/iqc and https://doi.org/10.17605/OSF.IO/AVD3Z for instructions and updates.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Clinical-Grade Validation of an Autofluorescence Virtual Staining System with Human Experts and a Deep Learning System for Prostate Cancer 97%
- Artificial Intelligence for Advance Requesting of Immunohistochemistry in Diagnostically Uncertain Prostate Biopsies 97%
- Tissue contamination challenges the credibility of machine learning models in real world digital pathology 96%
Similar papers in this journal
- Bladder Cancer Prognosis Using Deep Neural Networks and Histopathology Images 94%
- Independent assessment of a deep learning system for lymph node metastasis detection on the Augmented Reality Microscope 94%
- Development of an Interactive Web Dashboard to Facilitate the Reexamination of Pathology Reports for Instances of Underbilling of CPT Codes 94%
Similar papers in this journal
- Natural language inference for clinical registry curation 91%
- Quantification of abdominal fat from computed tomography using deep learning and its association with electronic health records in an academic biobank 91%
- ENRICHing Medical Imaging Training Sets Enables More Efficient Machine Learning 90%
Similar papers in this journal
- Weakly supervised learning for multi-organ adenocarcinoma classification in whole slide images 93%
- FalseColor-Python: a rapid intensity-leveling and digital-staining package for fluorescence-based slide-free digital pathology 93%
- Identifying Transcriptomic Correlates of Histology using Deep Learning 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.