Journal of Medical Imaging
● SPIE-Intl Soc Optical Eng
All preprints, ranked by how well they match Journal of Medical Imaging's content profile, based on 11 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Chang, S.; Wintergerst, G. A.; Carlson, C. J.; Yin, H.; Scarpato, K. R.; Luckenbaugh, A. N.; Chang, S.; Kolouri, S.; Bowden, A. K.
Show abstract
Bladder cancer is 10th most common malignancy and carries the highest treatment cost among all cancers. The high cost of bladder cancer treatment stems from its high recurrence rate, which necessitates frequent surveillance. White light cystoscopy (WLC), the standard of care surveillance tool to examine the bladder for lesions, has limited sensitivity for early-stage bladder cancer. Blue light cystoscopy (BLC) utilizes a fluorescent dye to induce contrast in cancerous regions, improving the sensitivity of detection by 43%. Nevertheless, the added cost and lengthy administration time of the dye limits the availability of BLC for surveillance. Here, we report the first demonstration of digital staining on clinical endoscopy videos collected with standard-of-care clinical equipment to convert WLC images to accurate BLC-like images. We introduce key pre-processing steps to circumvent color and brightness variations in clinical datasets needed for successful model performance; the results show excellent qualitative and quantitative agreement of the digitally stained WLC (dsWLC) images with ground truth BLC images as measured through staining accuracy analysis and color consistency assessment. In short, dsWLC can provide the fluorescent contrast needed to improve the detection sensitivity of bladder cancer, thereby increasing the accessibility of BLC contrast for bladder cancer surveillance use without the cost and time burden associated with the dye and specialized equipment.
Bonn, S.; Zimmermann, M.; Sauter, G.; Bengtsson, E.; Huber, T. B.; Baumbach, J.; Lennartz, M.; Fuhlert, P.; Witte, A.
Show abstract
BackgroundVision Foundation Models (VFM) have emerged as a promising approach for computational pathology, offering scalable feature representations that may reduce labelled-data requirements and improve robustness to variation in tissue preparation and digitisation. However, VFM decoder and dataset size requirements as well as the performance under real-world domain shifts remain unclear. MethodsWe evaluated six contemporary VFMs on a protocol-variant Prostate Cancer (PCa) dataset comprising 37 683 tissue microarray spot images from 10 412 patients. The dataset includes six controlled domain shifts arising from differences in staining duration, section thickness, scanner type, and sampling location. Two clinically relevant downstream tasks were examined: ISUP grading and 5-year relapse prediction. We compared two decoder architectures, quantified dataset-size requirements using a saturation analysis (45-5727 samples), and assessed cross-domain robustness using out-of-domain test sets. FindingsLarger VFMs consistently outperformed smaller models in peak accuracy and robustness metrics. Contrary to expectations of data efficiency, all models showed strong dependence on training-set size, requiring at least 1000 samples to approach stable results. All VFMs showed notable degradation under protocol-level domain shifts, with performance reductions of 4 to 13 percentage points in both cancer grading and relapse prediction, although larger models exhibited somewhat greater robustness. Furthermore, KNN-based probing performed substantially worse than a decoder-based approach across all architectures. InterpretationsDespite their strong representational capacity, current VFMs do not yet provide reliable domain generalisation or data-efficient performance in computational pathology. Decoder design remains essential, and substantial amounts of labelled data are still required to achieve clinically meaningful accuracy. Further advances in pre-training strategies, decoder architectures, and domain adaptation methods will be crucial for translating VFMs into robust clinical tools. Research in contextO_ST_ABSEvidence before this studyC_ST_ABSThis study focuses on pathology foundation models, which offer promising improvements in performance, data requirements, and robustness to domain shifts for computational pathology. To identify relevant studies, we searched in Google Scholar for research published before 1 April 2025. We searched for studies introducing novel foundation models trained on pathology images, or reviews comparing those models in terms of performance or robustness. The search terms used were computational pathology, benchmarking, review and foundation model, as well combinations of these terms. We found that many studies focus on increasing the complexity of pathology foundation models while using increasingly extensive and heterogeneous pre-training datasets. Various benchmarking studies demonstrate the superior performance and robustness of more recent and larger foundation models. However, these studies have limitations in their evaluation datasets. Either they cover only a domain shift due to a different scanner device, or they have small sample sizes. We also identified a research gap regarding the requirement for large datasets to train a decoder based on a pathology foundation model for a specific downstream task. Added value of this studyThe goal of this study was to evaluate the necessity of large downstream task datasets and the domain shift robustness of multiple pathology foundation models. For this purpose, we used our internal protocol-variant prostate cancer dataset, which provides a controlled evaluation setup as multiple domain shift types have been intentionally and separately introduced for different sub-datasets. Our saturation analysis revealed that at least 1000 samples were necessary to achieve good performance. Furthermore, our findings show that none of the evaluated foundation models are robust against all of our domain shifts, though larger models generally perform better. Implications of all the available evidenceThis study reveals that increasing the capacity of pathology foundation models improves performance and robustness. However, we demonstrated that all models exhibit some degree of performance degradation for certain domain shifts and require substantial datasets for training on downstream tasks. These limitations demonstrate that pathology foundation models do not fully address the issues of robustness and data requirements.
Hua, Y.; Mukkamala, A.; Estrada, C.; Li, M. L.; Wang, H.-H.
Show abstract
ObjectiveThe urinary tract dilation (UTD) classification system provides objective assessment relevant to hydronephrosis management for children. However, the lack of uniform language regarding UTD in radiology reports leads to significant difficulty in both clinical management and research. We seek to develop a unified multi-task/multi-class model that can effectively extract UTD components and classifications from early postnatal ultrasound (US) reports. MethodsRadiology records from our institution were reviewed to identify infants aged 0-90 days undergoing early ultrasound for antenatal UTD. The report and images were reviewed by the study team to create the ground truth of UTD classification and components (primary outcome). Bio_ClinicalBERT, a variant of the Bidirectional Encoder Representations from Transformers (BERT) model, was used as the embedding layers of the classification model. The model was fine-tuned with 11 linear classification layers. All but the last BERT layer were frozen during the fine-tuning process. The model performance was evaluated with five-fold cross-validation with an 80:20 train-test ratio. Results2460 early (0-90 days) US reports were included. The five-fold cross-validated model performance is satisfactory (Weighted F1 > 0.9 for all UTD components). We report the weighted F1 scores, accuracies, and standard deviations for all 11 tasks and their average performance. ConclusionsBy applying deep state-of-the-art NLP neural networks, we developed a high-performing, efficient, and scalable solution to extract UTD components from unstructured ultrasound reports using one single multi-task model. This can potentially help standardize and facilitate large-scale computer vision research for pediatric hydronephrosis. Key Words: machine learning, efficiency, ambulatory care, forecasting
Raad, R.; Ray, D.; Varghese, B.; Hwang, D.; Gill, I.; Duddalwar, V.; Oberai, A. A.
Show abstract
Image imputation refers to the task of generating a type of medical image given images of another type. This task becomes challenging when the difference between the available images, and the image to be imputed is large. In this manuscript, one such application, derived from the dynamic contrast enhanced computed tomography (CECT) imaging of the kidneys, is considered: given an incomplete sequence of three CECT images, we are required to the impute the missing image. This task is posed as one of probabilistic inference and a generative algorithm to generate samples of the imputed image, conditioned on the available images, is developed, trained, and tested. The output of this algorithm is the "best guess" of the imputed image, and a pixel-wise image of variance in the imputation. It is demonstrated that this best guess is more accurate than those generated by other, deterministic deep-learning based algorithms, including ones which utilize additional information and more complex loss terms. It is also shown the pixel-wise variance image, which quantifies the confidence in the reconstruction, can be used to determine whether the result of the imputation meets a specified accuracy threshold and is therefore appropriate for a downstream task.
Lutnick, B. R.; Sarder, P.
Show abstract
Segmentation of histology tissue whole side images is an important step for tissue analysis. Given enough annotated training data modern neural networks are capable accurate reproducible segmentation, however, the annotation of training datasets is time consuming. Techniques such as human in the loop annotation attempt to reduce this annotation burden, but still require a large amount of initial annotation. Semi-supervised learning, a technique which leverages both labeled and unlabeled data to learn features has shown promise for easing the burden of annotation. Towards this goal, we employ a recently published semi-supervised method: datasetGAN for the segmentation of glomeruli from renal biopsy images. We compare the performance of models trained using datasetGAN and traditional annotation and show that datasetGAN significantly reduces the amount of annotation required to develop a highly performing segmentation model. We also explore the usefulness of using datasetGAN for transfer learning and find that this greatly enhances the performance when a limited number of whole slide images are used for training.
Tariq, A.; Patel, B.; Banerjee, I.
Show abstract
Self-supervised pretraining can reduce the amount of labeled training data needed by pre-learning fundamental visual characteristics of the medical imaging data. In this study, we investigate several self-supervised training strategies for chest computed tomography exams and their effects of downstream applications. we bench-mark five well-known self-supervision strategies (masked image region prediction, next slice prediction, rotation prediction, flip prediction and denoising) on 15M chest CT slices collected from four sites of Mayo Clinic enterprise. These models were evaluated for two downstream tasks on public datasets; pulmonary embolism (PE) detection (classification) and lung nodule segmentation. Image embeddings generated by these models were also evaluated for prediction of patient age, race, and gender to study inherent biases in models understanding of chest CT exams. Use of pretraining weights, especially masked regions prediction based weights, improved performance and reduced computational effort needed for downstream tasks compared to task-specific state-of-the-art (SOTA) models. Performance improvement for PE detection was observed for training dataset sizes as large as [Formula] with maximum gain of 5% over SOTA. Segmentation model initialized with pretraining weights learned twice as fast as randomly initialized model. While gender and age predictors built using self-supervised training weights showed no performance improvement over randomly initialized predictors, the race predictor experienced a 10% performance boost when using self-supervised training weights. We released models and weights under open-source academic license. These models can then be finetuned with limited task-specific annotated data for a variety of downstream imaging tasks thus accelerating research in biomedical imaging informatics.
Castelo, A.; O'Connor, C.; Gupta, A. C.; Anderson, B. M.; Woodland, M.; Altaie, M.; Koay, E. J.; Odisio, B. C.; Tang, T. T.; Brock, K. K.
Show abstract
Artificial intelligence (AI) based segmentation has many medical applications but limited curated datasets challenge model training; this study compares the impact of dataset annotation quality and quantity on whole liver AI segmentation performance. We obtained 3,089 abdominal computed tomography scans with whole-liver contours from MD Anderson Cancer Center (MDA) and a MICCAI challenge. A total of 249 scans were withheld for testing of which 30, MICCAI challenge data, were reserved for external validation. The remaining scans were divided into mixed-curation and highly-curated groups, randomly sampled into sub-datasets of various sizes, and used to train 3D nnU-Net segmentation models. Dice similarity coefficients (DSC), surface DSC with 2mm margins (SD 2mm), the 95th percentile of Hausdorff distance (HD95), and 2D axial slice DSC (Slice DSC) were used to evaluate model performance. The highly curated, 244-scan model (DSC=0.971, SD 2mm=0.958, HD95=2.98mm) performed insignificantly different on 3D evaluation metrics to the mixed-curation 2,840-scan model (DSC=0.971 [p>.999], SD 2mm=0.958 [p>.999], HD95=2.87mm [p>.999]). The 710-scan mixed-curation (Slice DSC=0.929) significantly outperformed the highly curated, 244-scan model (Slice DSC=0.923 [p=0.012]) on the 30 external scans. Highly curated datasets yielded equivalent performance to datasets that were a full order of magnitude larger. The benefits of larger, mixed-curation datasets are evidenced in model generalizability metrics and local improvements. In conclusion, tradeoffs between dataset quality and quantity for model training are nuanced and goal dependent.
Grigis, A.; Alentorn, A.; Frouin, V.
Show abstract
Developing an automatic tumor detector for MRI medical images is a major challenge in neuro-oncology. The availability of such a tool would be a valuable assistance for the radiologists. Numerous works have tried to segment the tumor tissues, others have attempted to localize the tumor globally. In this work we focus on this second class of methods and we compare two drastically different strategies. The first one is an assumption-free anomaly detector build over a Variational Auto-Encoder (VAE), and the second one is a VGG classifier that embed Attention-Gated (AG) units to focus on the target structures at almost no additional computational cost. This comparison is first conducted on the publicly available BraTS glioma dataset for which published performance results can serve as reference, and extended as such (ie., without transfer learning) to two internal image datasets, namely Primary Central Nervous System Lymphoma (PCNSL) and Metastasis. The results demonstrate that the VAE and AG-VGG strategies can be used, up to a certain extent, to localize brain tumors.
Parikh, K.; Mathew, T. J.
Show abstract
With the growing amount of COVID-19 cases, especially in developing countries with limited medical resources, it is essential to accurately and efficiently diagnose COVID-19. Due to characteristic ground-glass opacities (GGOs) and other types of lesions being present in both COVID-19 and other acute lung diseases, misdiagnosis occurs often -- 26.6% of the time in manual interpretations of CT scans. Current deep-learning models can identify COVID-19 but cannot distinguish it from other common lung diseases like bacterial pneumonia. Concretely, COVision is a deep-learning model that can differentiate COVID-19 from other common lung diseases, with high specificity using CT scans and other clinical factors. COVision was designed to minimize overfitting and complexity by decreasing the number of hidden layers and trainable parameters while still achieving superior performance. Our model consists of two parts: the CNN which analyzes CT scans and the CFNN (clinical factors neural network) which analyzes clinical factors such as age, gender, etc. Using federated averaging, we ensembled our CNN with the CFNN to create a comprehensive diagnostic tool. After training, our CNN achieved an accuracy of 95.8% and our CFNN achieved an accuracy of 88.75% on a validation set. We found a statistical significance that COVision performs better than three independent radiologists with at least 10 years of experience, especially in differentiating COVID-19 from pneumonia. We analyzed our CNNs activation maps through Grad-CAMs and found that lesions in COVID-19 presented peripherally, closer to the pleura, whereas pneumonia lesions presented centrally.
Van Booven, D. J.; Chen, C.-B.; Kryvenko, O.; Punnen, S.; Sandoval, V.; Malpani, S.; Noman, A.; Ismael, F.; Briseno, A.; Wang, Y.; Arora, H.
Show abstract
Prostate cancer (PCa) poses significant challenges for timely diagnosis and prognosis, leading to high mortality rates and increased disease risk and treatment costs. Recent advancements in machine learning and digital imagery offer promising potential for developing automated and objective assessment pipelines that can reduce human capital and resource costs. However, the reliance of AI models on large amounts of clinical data for training presents a significant challenge, as this data is often biased, lacking diversity, and not readily available. Here we aim to address this limitation by employing customized generative adversarial network (GAN) models to produce high-quality synthetic images of different PCa grades (radical prostatectomy (RP)) and needle biopsies, which were customized to account for the granularity associated with each Gleason grade. The generated images were subjected to multiple rounds of benchmarking, quantifications and quality control assessment before being used to train an AI model (EfficientNet) for grading digital histology images of adenocarcinoma specimens (RP sections) and needle biopsies obtained from the PANDA challenge repository. Validation was performed using the AI model trained with synthetic data to grade digital histology from the cancer genome atlas (TCGA) (RP sections) and needle biopsy data from Radboud University Medical Center and Karolinska Institute. Results demonstrated that the AI model trained with a combination of image patches derived from original and enhanced synthetic images outperformed the model trained with original digital histology images. Together, this study demonstrates the potential of customized GAN models to generate a large cohort of synthetic data that can train AI models to effectively grade PCa specimens. This approach could potentially eliminate the need for extensive clinical data for training any AI model in the domain of digital imagery, leading to cost and time-effective diagnosis and prognosis.
Wan, S.-Y.; Chen, W.-Y.
Show abstract
Accurate segmentation of nasal and paranasal sinus structures from CT scans is critical for surgical planning and treatment evaluation in rhinology. However, the complex anatomical topology and thin-wall boundaries of these structures pose significant challenges for automated segmentation methods. We propose AFS-DSN (Adaptive Frequency-Spatial Dual-Stream Network), a novel deep learning architecture that integrates multi-scale wavelet decomposition with spatial feature learning for binary segmentation of the nasal cavity complex. Our method employs a dual-stream encoder with frequency branch utilizing three wavelet scales (db1, db2, db4) to capture 24 frequency sub-bands, enabling enhanced boundary detection in anatomically challenging regions. Cross-domain attention and adaptive routing mechanisms dynamically fuse spatial and frequency features based on local tissue characteristics. We formulate the task as binary segmentation where all five anatomical structures (maxillary sinus, sphenoid sinus, ethmoid sinus, frontal sinus, and nasal cavity) are treated as a unified foreground region against the background, prioritizing clinical boundary detection over individual structure differentiation. Evaluated on the NasalSeg dataset (130 CT volumes) with a 70/15/15 train/validation/test split, AFS-DSN achieves 94.34% {+/-} 2.30% overall Dice coefficient with statistically significant improvements in thin-wall regions (91.34% vs. 90.57% baseline, p=0.004) and statistically significant improvement in Surface Dice at 1mm tolerance (0.874 vs. 0.868 baseline, p=0.010), demonstrating enhanced boundary precision while maintaining sub-second inference time, making the method suitable for surgical planning applications where sub-millimeter accuracy is clinically relevant. To address concerns regarding model complexity, we further introduce AFS-DSN-Lite, a parameter-efficient variant (27.41M parameters) that achieves comparable performance (94.37% Dice) through depthwise separable convolutions, and validate robustness via 3-fold cross-validation (mean Dice: 94.59% {+/-} 0.31%).
Hari, S. N.; Nyman, J.; Mehta, N.; Jiang, B.; Rosenthal, J.; Sengupta, E.; Dietlein, F.; Umeton, R.; Van Allen, E. M.
Show abstract
Computer vision (CV) approaches applied to digital pathology have informed biological discovery and development of tools to help inform clinical decision-making. However, batch effects in the images have the potential to introduce spurious confounders and represent a major challenge to effective analysis and interpretation of these data. Standard methods to circumvent learning such confounders include (i) application of image augmentation techniques and (ii) examination of the learning process by evaluating through external validation (e.g., unseen data coming from a comparable dataset collected at another hospital). Here, we show that the source site of a histopathology slide can be learned from the image using CV algorithms in spite of image augmentation, and we explore these source site predictions using interpretability tools. A CV model trained using Empirical Risk Minimization (ERM) risks learning this source-site signal as a spurious correlate in the weak-label regime, which we abate by using a training method with abstention. We find that a patch based classifier trained using abstention outperformed a model trained using ERM by 9.9, 10 and 19.4% F1 in the binary classification tasks of identifying tumor versus normal tissue in lung adenocarcinoma, Gleason score in prostate adenocarcinoma, and tumor tissue grade in clear cell renal cell carcinoma, respectively, at the expense of up to 80% coverage (defined as the percent of tiles not abstained on by the model). Further, by examining the areas abstained by the model, we find that the model trained using abstention is more robust to heterogeneity, artifacts and spurious correlates in the tissue. Thus, a method trained with abstention may offer novel insights into relevant areas of the tissue contributing to a particular phenotype. Together, we suggest using data augmentation methods that help mitigate a digital pathology models reliance on potentially spurious visual features, as well as selecting models that can identify features truly relevant for translational discovery and clinical decision support.
Xia, Y.; Yu, Q.; Chu, L.; Kawamoto, S.; Park, S.; Liu, F.; Chen, J.; Zhu, Z.; Li, B.; Zhou, Z.; Lu, Y.; Wang, Y.; Shen, W.; Xie, L.; Zhou, Y.; Wolfgang, C.; Javed, A.; Fouladi, D. F.; Shayesteh, S.; Graves, J.; Blanco, A.; Zinreich, E. S.; Kinny-Koster, B.; Kinzler, K.; Hruban, R. H.; Vogelstein, B.; Yuille, A. L.; Fishman, E. K.
Show abstract
Tens of millions of abdominal images are obtained with computed tomography (CT) in the U.S. each year but pancreatic cancers are sometimes not initially detected in these images. We here describe a suite of algorithms (named FELIX) that can recognize pancreatic lesions from CT images without human input. Using FELIX, >95% of patients with pancreatic ductal adenocarcinomas were detected at a specificity of >95% in patients without pancreatic disease. FELIX may be able to assist radiologists in identifying pancreatic cancers earlier, when surgery and other treatments offer more hope for long-term survival.
Avesta, A. E.; Hossain, S.; Aboian, M.; Krumholz, H.; Aneja, S.
Show abstract
When an auto-segmentation model needs to be applied to a new segmentation task, multiple decisions should be made about the pre-processing steps and training hyperparameters. These decisions are cumbersome and require a high level of expertise. To remedy this problem, I developed self-configuring CapsNets (scCapsNets) that can scan the training data as well as the computational resources that are available, and then self-configure most of their design options. In this study, we developed a self-configuring capsule network that can configure its design options with minimal user input. We showed that our self-configuring capsule netwrok can segment brain tumor components, namely edema and enhancing core of brain tumors, with high accuracy. Out model outperforms UNet-based models in the absence of data augmentation, is faster to train, and is computationally more efficient compared to UNet-based models.
Mensah, S.; Atsu, E. K. A.; Ammah, P. N. T.
Show abstract
Brain tumors are one of the most life-threatening diseases, requiring precise and timely detection for effective treatment. Traditional methods for brain tumor detection rely heavily on manual analysis of MRI scans, which is time-consuming, subjective, and prone to human error. With advancements in deep learning, Convolutional Neural Networks (CNNs) have become popular for medical image analysis. However, CNNs are limited in their ability to capture spatial hierarchies and pose variations, which reduces their accuracy, particularly for tasks like brain tumor segmentation where precise spatial relationships are crucial. This research introduces a hybrid Capsule Neural Network (CapsNet) and ResNet50 model designed to overcome the limitations of traditional CNNs by capturing both spatial and pose information in MRI scans. The proposed model leverages ResNet50 for feature extraction and CapsNet for handling spatial relationships, leading to more accurate segmentation. The study evaluates the model on the BraTS2020 dataset and compares its performance to state-of-the-art CNN architectures, including U-Net and pure CNN models. The hybrid model, featuring a custom 5-cycle dynamic routing algorithm to enhance capsule agreement for tumor boundaries, achieved 98% accuracy and an F1-score of 0.87, demonstrating superior performance in detecting and segmenting brain tumors. This study pioneers the systematic evaluation of the ResNet50 + CapsNet hybrid on the BraTS2020 dataset, with a tailored class weighting scheme addressing class imbalance, improving effectiveness in identifying irregularly shaped tumors and smaller regions in identifying irregularly shaped tumors and smaller tumor regions. The study offers a robust solution for automating brain tumor detection. Future work will explore the use of Capsule Networks alone for brain tumor detection in MRI data and investigate alternative Capsule Network architectures, as well as their integration into clinical decision support systems.
Sivakumar, E.; Anand, A.
Show abstract
Computer vision and deep learning techniques, including convolutional neural networks (CNNs) and transformers, have increased the performance of medical image classification systems. However, training deep learning models using medical images is a challenging task that necessitates a substantial amount of annotated data. In this paper, we implement data augmentation strategies to tackle dataset imbalance in the VinDr-SpineXR dataset, which has a lower number of spine abnormality X-ray images compared to normal spine X-ray images. Geometric transformations and synthetic image generation using Generative Adversarial Networks are explored and applied to the abnormal classes of the dataset, and classifier performance is validated using VGG-16 and InceptionNet to identify the most effective augmentation technique. Additionally, we introduce a hybrid augmentation technique that addresses class imbalance, reduces computational overhead relative to a GAN-only approach, and achieves [~]99% validation accuracy with both classifiers across all three case studies.
Sassa, N.; Kameya, Y.; Takahashi, T.; Matsukawa, Y.; Majima, T.; Tsuruta, K.; Kobayashi, I.; Kajikawa, K.; Kawanishi, H.; Kurosu, H.; Yamagiwa, S.; Takahashi, M.; Hotta, K.; Yamada, K.; Yamamoto, T.
Show abstract
ObjectivesTo elucidate if synthetic contrast enhanced computed tomography (CECT) images created from plain CT images using deep neural networks (DNN) could be used for screening, clinical diagnosis, and postoperative follow-up of small-diameter renal tumors by comparing the concordance rate between real and synthetic CECT images and the diagnoses according to 10 urologists. MethodsThis retrospective, multicenter study included 155 patients (artificial intelligence training cohort [n=99], validation cohort [n=56]) who underwent surgery for small-diameter ([≤]40 mm) renal tumors, with the pathological diagnosis of renal cell carcinoma, during 2010-2020. Preoperatively, dynamic plain CT and CECT images were obtained. We created a learned DNN using pix2pix. We examined the quality of the synthetic CECT images created using this DNN and compared them with real CECT images using the zero-mean normalized cross-correlation parameter. We assessed concordance rates between real and synthetic images and diagnoses according to 10 urologists by creating a receiver operating characteristic curve and calculating the area under the curve (AUC). ResultsThe synthetic CECT images were highly concordant with the real CECT images, regardless of the existence or morphology of the renal tumor. Regarding the concordance rate, a greater AUC was obtained with synthetic CECT (AUC=0.892) than with only CT (AUC=0.720; p<0.001). ConclusionsThis study is the first to use DNN to create a high-quality synthetic CECT image that was highly concordant with a real CECT image. Synthetic CECT images could be used for urological diagnoses and clinical screening.
Bennett, J.; Woodland, M.; Castelo, A.; Altaie, M.; Antony, A.; Siddiqi, N. S.; Long, J. P.; Brock, K. K.
Show abstract
Deep learning models deployed in clinical imaging frequently encounter distribution shifts, yet most out-of-distribution (OOD) detection methods are evaluated only on controlled research datasets. As a result, it is unclear whether existing approaches can reliably identify segmentation failures that arise in real-world clinical practice. We evaluated six OOD detection methods on a deployed liver CT segmentation model (3D nnU-Net) using internal data from 400 patients and external data from 100 patients collected across nearly 70 sites in 7 countries. One method was Pairwise Surface DSC, a surface-based extension of Pairwise DSC, that we introduced. OOD performance was measured using sensitivity, AUROC, and balanced accuracy, with thresholds determined on an independent cohort of 400 patients using the Youden J statistic. Statistical significance was assessed using McNemar tests and stratified bootstraps ( = 0.05) with Benjamini-Hochberg correction. Pairwise Surface DSC was the top-performing method, with perfect sensitivities (1.00), near-perfect AUROCs (0.97 internal; 1.00 external), and the highest balanced accuracies (0.94 internal; 0.88 external; p<0.001). These results show that automated failure detection for liver CT segmentation is clinically feasible and that Pairwise Surface DSC is a promising candidate for deployment. Our code is available at https://github.com/mckellwoodland/liver_ct_ood_translation.
Levy, J. J.; Jackson, C. R.; Haudenschild, C. C.; Christensen, B. C.; Vaickus, L. J.
Show abstract
Image registration involves finding the best alignment between different images of the same object. In these tasks, the object in question is viewed differently in each of the images (e.g. different rotation or light conditions, etc.). In digital pathology, image registration aligns correspondent regions of tissue from different stereotactic viewpoints (e.g. subsequent deeper sections of the same tissue). These comparisons are important for histological analysis and can facilitate previously unavailable manipulations, such as 3D tissue reconstruction and cell-level alignment of immunohistochemical (IHC) and special stains. Several benchmarks have been established for evaluating image registration techniques for histological tissue; however, little work has evaluated the impact of scaling registration techniques to Giga-Pixel Whole Slide Images (WSI), which are large enough for significant memory limitations, and contain recurrent patterns and deformations that hinder traditional alignment algorithms. Furthermore, as tissue sections often contain multiple, discrete, smaller tissue fragments, it is unnecessary to align an entire image when the bulk of the image is background whitespace and tissue fragments orientations are often agnostic of each other. We present a methodology for circumventing large-scale image registration issues in histopathology and accompanying software. By removing background pixels, parsing the slide into discrete tissue segments, and matching, orienting and registering smaller segment pairs, we recovered registrations with lower Target Registration Error (TRE) when compared to utilizing the unmanipulated WSI. We tested our technique by having a pathologist annotate landmarks from 13 pairs of differently stained liver biopsy slides, performing WSI and segment-based registration techniques, and comparing overall TRE. Preliminary results demonstrate superior performance of registering segment pairs versus registering WSI (difference of median TRE of 44 pixels, p<0.001). Segment matching within WSI is an effective solution for histology image registration but requires further testing and validation to ensure its viability for stain translation and 3D histology analysis.
Sahin, S.; Diaz, E.; Rajagopal, A.; Abtahi, M.; Jones, S.; Dai, Q.; Kramer, S.; Wang, Z.; Larson, P. E. Z.
Show abstract
Current standard of care imaging practices cannot reliably differentiate among certain renal tumors such as benign oncocytoma and clear cell renal cell carcinoma (RCC), and between low and high grade RCCs. Previous work has explored using deep learning, radiomics, and texture analysis to predict renal tumor subtypes and differentiate between low and high grade RCCs with mixed success. To further this work, large diverse datasets are needed to improve model performance and provide strong evaluation sets. In this work, a dataset of 831 multi-phase 3D CT exams was curated. Each exam contains up to three contrast-enhanced CT phases. Tumor outlines or bounding boxes were annotated and registered to the image volumes. The pathology results for each tumor and relevant patient metadata are also included.