Back

Medical Physics

Wiley

All preprints, ranked by how well they match Medical Physics's content profile, based on 14 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Heart-centered positioning and tailored beam-shaping filtration for reduced radiation dose in coronary artery calcium imaging: a MESA study

Colvert, B.; Rigolli, M.; Craine, A.; Criqui, M.; Contijoch, F.

2021-07-22 radiology and imaging 10.1101/2021.07.18.21259666 medRxiv
Top 0.1%
34.6%
Show abstract

PurposeCardiac CT has a clear clinical role in the evaluation of coronary artery disease and assessment of coronary artery calcium (CAC) but the use of ionizing radiation limits clinical use. Beam shaping "bow-tie" filters determine the radiation dose and the effective scan field-of-view diameter (SFOV) by delivering higher X-ray fluence to a region centered at the isocenter. A method for positioning the heart near the isocenter could enable reduced SFOV imaging and reduce dose in cardiac scans. However, a predictive approach to center the heart, the extent to which heart centering can reduce the SFOV, and the associated dose reductions have not been assessed. The purpose of this study is to build a heart-centered patient positioning model, to test whether it reduces the SFOV required for accurate CAC scoring, and to quantify the associated reduction in radiation dose. MethodsThe location of 38,184 calcium lesions (3,151 studies) in the Multi-Ethnic Study of Atherosclerosis (MESA) were utilized to build a predictive heart-centered positioning model and compare the impact of SFOV on CAC scoring accuracy in heart-centered and conventional body-centered scanning. Then, the positioning model was applied retrospectively to an independent, contemporary cohort of 118 individuals (81 with CAC>0) at our institution to validate the models ability to maintain CAC accuracy while reducing the SFOV. In these patients, the reduction in dose associated with a reduced SFOV beam-shaping filter was quantified. ResultsHeart centering reduced the SFOV diameter 25.7% relative to body centering while maintaining high CAC scoring accuracy (0.82% risk reclassification rate). In our validation cohort, imaging at this reduced SFOV with heart-centered positioning and tailored beam-shaping filtration led to a 26.9% median dose reduction (25-75th percentile: 21.6 to 29.8%) without any calcium risk reclassification. ConclusionsHeart-centered patient positioning enables a significant radiation dose reduction while maintaining CAC accuracy.

2
Organ Finder a new AI-based organ segmentation tool for CT

Edenbrandt, L.; Enqvist, O.; Larsson, M.; Ulen, J.

2022-11-18 radiology and imaging 10.1101/2022.11.15.22282357 medRxiv
Top 0.1%
31.6%
Show abstract

BackgroundAutomated organ segmentation in computed tomography (CT) is a vital component in many artificial intelligence-based tools in medical imaging. This study presents a new organ segmentation tool called Organ Finder 2.0. In contrast to most existing methods, Organ Finder was trained and evaluated on a rich multi-origin dataset with both contrast and non-contrast studies from different vendors and patient populations. ApproachA total of 1,171 CT studies from seven different publicly available CT databases were retrospectively included. Twenty CT studies were used as test set and the remaining 1,151 were used to train a convolutional neural network. Twenty-two different organs were studied. Professional annotators segmented a total of 5,826 organs and segmentation quality was assured manually for each of these organs. ResultsOrgan Finder showed high agreement with manual segmentations in the test set. The average Dice index over all organs was 0.93 and the same high performance was found for four different subgroups of the test set based on the presence or absence of intravenous and oral contrast. ConclusionsAn AI-based tool can be used to accurately segment organs in both contrast and non-contrast CT studies. The results indicate that a large training set and high-quality manual segmentations should be used to handle common variations in the appearance of clinical CT studies.

3
Clinical Evaluation of a Novel Deep Learning-Based Auto-Segmentation Software: Utility and Potential Pitfalls

Tozuka, R.; Saito, M.; Matsuda, M.; Akita, T.; Nemoto, H.; Komiyama, T.; Kadoya, N.; Jingu, K.; Onishi, H.

2026-01-11 radiology and imaging 10.64898/2026.01.08.26343652 medRxiv
Top 0.1%
31.1%
Show abstract

BackgroundAccurate contouring of target volumes and organs at risk is critical for radiotherapy. While deep learning (DL) models offer efficient automation, their generalizability to real-world clinical cases containing anatomical variations and artifacts requires rigorous validation. PurposeTo evaluate the clinical accuracy and robustness of RatoGuide, a novel DL-based auto-segmentation software, using a dataset derived from routine clinical practice including atypical cases. MethodsThis single-center retrospective study included 36 patients treated for head and neck, thoracic, abdominal, and pelvic cancers. The cohort was intentionally selected to encompass diverse anatomies and artifacts (e.g., pacemakers, artificial femoral head replacement). Auto-contours generated by RatoGuide were compared with expert-approved manual contours. Performance was evaluated quantitatively using the Dice Similarity Coefficient (DSC) and 95th percentile Hausdorff Distance (HD95), and qualitatively via a 5-point visual assessment scale (higher is better) by four independent reviewers. A score of [≤]2 by multiple reviewers was defined as failure. ResultsOverall, the mean DSC, HD95, and visual assessment score were 0.79 {+/-} 0.19, 6.35 {+/-} 12.2 mm, and 3.65 {+/-} 0.88, respectively. The mean DSC exceeded 0.8 in 62% (23/37 organ structures) of the evaluated structure types, and a total of 93.5% (315/337) of all contours were considered clinically acceptable based on visual evaluation . However, lower performance was observed in small structures (e.g., optic chiasm) and low-contrast organs (e.g., esophagus). ConclusionsRatoGuide demonstrated favorable performance for major organs across various anatomical regions, consistent with benchmarks reported in the literature. However, performance variability in atypical cases underscores the necessity of rigorous visual verification by experts for clinical implementation.

4
Convolutional Encoder-Decoder Networks for Volumetric Computed Tomography Surviews from Single- and Dual-View Topograms

Shapira, N.; Bharthulwar, S.; Noel, P. B.

2022-05-21 radiology and imaging 10.1101/2022.05.17.22275229 medRxiv
Top 0.1%
31.1%
Show abstract

Computed tomography (CT) is an extensively used imaging modality capable of generating detailed images of a patients internal anatomy for diagnostic and interventional procedures. High-resolution volumes are created by measuring and combining information along many radiographic projection angles. In current medical practice, single and dual-view two-dimensional (2D) topograms are utilized for planning the proceeding diagnostic scans and for selecting favorable acquisition parameters, either manually or automatically, as well as for dose modulation calculations. In this study, we develop modified 2D to three-dimensional (3D) encoder-decoder neural network architectures to generate CT-like volumes from single and dual-view topograms. We validate the developed neural networks on synthesized topograms from publicly available thoracic CT datasets. Finally, we assess the viability of the proposed transformational encoder-decoder architecture on both common image similarity metrics and quantitative clinical use case metrics, a first for 2D-to-3D CT reconstruction research. According to our findings, both single-input and dual-input neural networks are able to provide accurate volumetric anatomical estimates. The proposed technology will allow for improved (i) planning of diagnostic CT acquisitions, (ii) input for various dose modulation techniques, and (iii) recommendations for acquisition parameters and/or automatic parameter selection. It may also provide for an accurate attenuation correction map for positron emission tomography (PET) with only a small fraction of the radiation dose utilized.

5
Application of simultaneous uncertainty quantification for image segmentation with probabilistic deep learning: Performance benchmarking of oropharyngeal cancer target delineation as a use-case

Sahlsten, J.; Jaskari, J.; Wahid, K. A.; Ahmed, S.; Glerean, E.; He, R.; Kann, B.; Makitie, A. A.; Fuller, C. D.; Naser, M. A.; Kaski, K.

2023-02-24 radiology and imaging 10.1101/2023.02.20.23286188 medRxiv
Top 0.1%
30.9%
Show abstract

BackgroundOropharyngeal cancer (OPC) is a widespread disease, with radiotherapy being a core treatment modality. Manual segmentation of the primary gross tumor volume (GTVp) is currently employed for OPC radiotherapy planning, but is subject to significant interobserver variability. Deep learning (DL) approaches have shown promise in automating GTVp segmentation, but comparative (auto)confidence metrics of these models predictions has not been well-explored. Quantifying instance-specific DL model uncertainty is crucial to improving clinician trust and facilitating broad clinical implementation. Therefore, in this study, probabilistic DL models for GTVp auto-segmentation were developed using large-scale PET/CT datasets, and various uncertainty auto-estimation methods were systematically investigated and benchmarked. MethodsWe utilized the publicly available 2021 HECKTOR Challenge training dataset with 224 co-registered PET/CT scans of OPC patients with corresponding GTVp segmentations as a development set. A separate set of 67 co-registered PET/CT scans of OPC patients with corresponding GTVp segmentations was used for external validation. Two approximate Bayesian deep learning methods, the MC Dropout Ensemble and Deep Ensemble, both with five submodels, were evaluated for GTVp segmentation and uncertainty performance. The segmentation performance was evaluated using the volumetric Dice similarity coefficient (DSC), mean surface distance (MSD), and Hausdorff distance at 95% (95HD). The uncertainty was evaluated using four measures from literature: coefficient of variation (CV), structure expected entropy, structure predictive entropy, and structure mutual information, and additionally with our novel Dice-risk measure. The utility of uncertainty information was evaluated with the accuracy of uncertainty-based segmentation performance prediction using the Accuracy vs Uncertainty (AvU) metric, and by examining the linear correlation between uncertainty estimates and DSC. In addition, batch-based and instance-based referral processes were examined, where the patients with high uncertainty were rejected from the set. In the batch referral process, the area under the referral curve with DSC (R-DSC AUC) was used for evaluation, whereas in the instance referral process, the DSC at various uncertainty thresholds were examined. ResultsBoth models behaved similarly in terms of the segmentation performance and uncertainty estimation. Specifically, the MC Dropout Ensemble had 0.776 DSC, 1.703 mm MSD, and 5.385 mm 95HD. The Deep Ensemble had 0.767 DSC, 1.717 mm MSD, and 5.477 mm 95HD. The uncertainty measure with the highest DSC correlation was structure predictive entropy with correlation coefficients of 0.699 and 0.692 for the MC Dropout Ensemble and the Deep Ensemble, respectively. The highest AvU value was 0.866 for both models. The best performing uncertainty measure for both models was the CV which had R-DSC AUC of 0.783 and 0.782 for the MC Dropout Ensemble and Deep Ensemble, respectively. With referring patients based on uncertainty thresholds from 0.85 validation DSC for all uncertainty measures, on average the DSC improved from the full dataset by 4.7% and 5.0% while referring 21.8% and 22% patients for MC Dropout Ensemble and Deep Ensemble, respectively. ConclusionWe found that many of the investigated methods provide overall similar but distinct utility in terms of predicting segmentation quality and referral performance. These findings are a critical first-step towards more widespread implementation of uncertainty quantification in OPC GTVp segmentation.

6
The Effect of Image Resolution on the Performance of Deep Learning Algorithms in Detecting Calcaneus Fractures on X-Ray

Yee, N. J.; Taseh, A.; Ghandour, S.; Sirls, E.; Halai, M.; Whyne, C.; DiGiovanni, C. W.; Kwon, J. Y.; Ashkani-Esfahani, S. J.

2025-09-07 orthopedics 10.1101/2025.09.04.25334786 medRxiv
Top 0.1%
29.2%
Show abstract

PurposeTo evaluate convolutional neural network (CNN) model training strategies that optimize the performance of calcaneus fracture detection on radiographs at different image resolutions. Materials and MethodsThis retrospective study included foot radiographs from a single hospital between 2015 and 2022 for a total of 1,775 x-ray series (551 fractures; 1,224 without) and was split into training (70%), validation (15%), and testing (15%). ImageNet pre-trained ResNet models were fine-tuned on the dataset. Three training strategies were evaluated: 1) single size: trained exclusively on 128x128, 256x256, 512x512, 640x640, or 900x900 radiographs (5 model sets); 2) curriculum learning: trained exclusively on 128x128 radiographs then exclusively on 256x256, then 512x512, then 640x640, and finally on 900x900 (5 model sets); and 3) multi-scale augmentation: trained on x-ray images resized along continuous dimensions between 128x128 to 900x900 (1 model set). Inference time and training time were compared. ResultsMulti-scale augmentation trained models achieved the highest average area under the Receiver Operating Characteristic curve of 0.938 [95% CI: 0.936 - 0.939] for a single model across image resolutions compared to the other strategies without prolonging training or inference time. Using the optimal model sets, curriculum learning had the highest sensitivity on in-distribution low-resolution images (85.4% to 90.1%) and on out-of-distribution high-resolution images (78.2% to 89.2%). However, curriculum learning models took significantly longer to train (11.8 [IQR: 11.1-16.4] hours; P<.001). ConclusioWhile 512x512 images worked well for fracture identification, curriculum learning and multi-scale augmentation training strategies algorithmically improved model robustness towards different image resolutions without requiring additional annotated data. Summary statementDifferent deep learning training strategies affect performance in detecting calcaneus fractures on radiographs across in- and out-of-distribution image resolutions, with a multi-scale augmentation strategy conferring the greatest overall performance improvement in a single model. Key pointsO_LITraining strategies addressing differences in radiograph image resolution (or pixel dimensions) could improve deep learning performance. C_LIO_LIThe highest average performance across different image resolutions in a single model was achieved by multi-scale augmentation, where the sampled training dataset is uniformly resized between square resolutions of 128x128 to 900x900. C_LIO_LICompared to model training on a single image resolution, sequentially training on increasingly higher resolution images up to 900x900 (i.e., curriculum learning) resulted in higher fracture detection performance on images resolutions between 128x128 and 2048x2048. C_LI

7
TCIA Radiology Image Processing for AI and Radiomics

Rich, J. M.; Kang, R.; Jin, D.; Subramanian, S.; Duddalwar, V.; Pachter, L.

2026-06-24 radiology and imaging 10.64898/2026.06.15.26354651 medRxiv
Top 0.1%
28.4%
Show abstract

We developed a standardized, reproducible preprocessing framework for computed tomography (CT) imaging data from multi-institutional repositories such The Cancer Imaging Archive (TCIA), enabling consistent radiomics and artificial intelligence (AI) analyses. Imaging data from TCGA-KIRC patients available on TCIA were used as a representative heterogeneous dataset characterized by variation in acquisition protocols, inconsistent metadata, and differing image quality. The proposed modular pipeline includes series filtering, DICOM-to-NIfTI conversion, orientation harmonization to a canonical coordinate system, voxel spacing normalization, intensity clipping and normalization, segmentation integration, and metadata validation, and is implemented in a reproducible, notebook-based framework compatible with common radiomics and deep learning workflows. This pipeline standardizes imaging data into analysis-ready volumes with consistent geometry, intensity distributions, and spatial alignment, reducing non-biological variability that can adversely affect radiomic feature stability and model performance. The modular design enables task-specific adaptation of individual preprocessing steps while maintaining overall consistency. Although demonstrated on TCIA, this framework is generalizable to other heterogeneous imaging datasets and provides a foundation for robust, large-scale computational imaging studies.

8
Dual-Branch Efficient Net Architecture for ACL Tear Detection in Knee MRI

kota, T.; Garofalaki, K.; Whitely, F.; Evdokimenko, E.; Smartt, E.

2025-09-13 orthopedics 10.1101/2025.09.09.25335391 medRxiv
Top 0.1%
26.9%
Show abstract

We propose a deep learning approach for detecting anterior cruciate ligament (ACL) tears from knee MRI using a dual-branch convolutional architecture. The model independently processes sagittal and coronal MRI sequences using EfficientNet-B2 backbones with spatial attention modules, followed by a late fusion classifier for binary prediction. MRI volumes are standardized to a fixed number of slices, and domain-specific normalization and data augmentation are applied to enhance model robustness. Trained on a stratified 80/20 split of the MRNet dataset, our best model--using the Adam optimizer and a learning rate of 1e-4--achieved a validation AUC of 0.98 and a test AUC of 0.93. These results show strong predictive performance while maintaining computational efficiency. This work demonstrates that accurate diagnosis is achievable using only two anatomical planes and sets the stage for further improvements through architectural enhancements and broader data integration.

9
Investigation of Autosegmentation Techniques on T2-Weighted MRI for Off-line Dose Reconstruction in MR-Linac Adapt to Position Workflow for Head and Neck Cancers

McDonald, B. A.; Cardenas, C.; O'Connell, N.; Ahmed, S.; Naser, M. A.; Wahid, K. A.; Xu, J.; Thill, D.; Zuhour, R.; Mesko, S.; Augustyn, A.; Buszek, S. M.; Grant, S.; Chapman, B. V.; Bagley, A.; He, R.; Mohamed, A. S. R.; Christodouleas, J. P.; Brock, K. K.; Fuller, C. D.

2021-10-01 radiology and imaging 10.1101/2021.09.30.21264327 medRxiv
Top 0.1%
26.6%
Show abstract

PurposeIn order to accurately accumulate delivered dose for head and neck cancer patients treated with the Adapt to Position workflow on the 1.5T magnetic resonance imaging (MRI)-linear accelerator (MR-linac), the low-resolution T2-weighted MRIs used for daily setup must be segmented to enable reconstruction of the delivered dose at each fraction. In this study, our goal is to evaluate various autosegmentation methods for head and neck organs at risk (OARs) on on-board setup MRIs from the MR-linac for off-line reconstruction of delivered dose. MethodsSeven OARs (parotid glands, submandibular glands, mandible, spinal cord, and brainstem) were contoured on 43 images by seven observers each. Ground truth contours were generated using a simultaneous truth and performance level estimation (STAPLE) algorithm. 20 autosegmentation methods were evaluated in ADMIRE: 1-9) atlas-based autosegmentation using a population atlas library (PAL) of 5/10/15 patients with STAPLE, patch fusion (PF), random forest (RF) for label fusion; 10-19) autosegmentation using images from a patients 1-4 prior fractions (individualized patient prior (IPP)) using STAPLE/PF/RF; 20) deep learning (DL) (3D ResUNet trained on 43 ground truth structure sets plus 45 contoured by one observer). Execution time was measured for each method. Autosegmented structures were compared to ground truth structures using the Dice similarity coefficient, mean surface distance, Hausdorff distance, and Jaccard index. For each metric and OAR, performance was compared to the inter-observer variability using Dunns test with control. Methods were compared pairwise using the Steel-Dwass test for each metric pooled across all OARs. Further dosimetric analysis was performed on three high-performing autosegmentation methods (DL, IPP with RF and 4 fractions (IPP_RF_4), IPP with 1 fraction (IPP_1)), and one low-performing (PAL with STAPLE and 5 atlases (PAL_ST_5)). For five patients, delivered doses from clinical plans were recalculated on setup images with ground truth and autosegmented structure sets. Differences in maximum and mean dose to each structure between the ground truth and autosegmented structures were calculated and correlated with geometric metrics. ResultsDL and IPP methods performed best overall, all significantly outperforming inter-observer variability and with no significant difference between methods in pairwise comparison. PAL methods performed worst overall; most were not significantly different from the inter-observer variability or from each other. DL was the fastest method (33 seconds per case) and PAL methods the slowest (3.7 - 13.8 minutes per case). Execution time increased with number of prior fractions/atlases for IPP and PAL. For DL, IPP_1, and IPP_RF_4, the majority (95%) of dose differences were within {+/-}250 cGy from ground truth, but outlier differences up to 785 cGy occurred. Dose differences were much higher for PAL_ST_5, with outlier differences up to 1920 cGy. Dose differences showed weak but significant correlations with all geometric metrics (R2 between 0.030 and 0.314). ConclusionsThe autosegmentation methods offering the best combination of performance and execution time are DL and IPP_1. Dose reconstruction on on-board T2-weighted MRIs is feasible with autosegmented structures with minimal dosimetric variation from ground truth, but contours should be visually inspected prior to dose reconstruction in an end-to-end dose accumulation workflow.

10
Automated Segmentation of Head and Neck Cancer from CT Images Using 3D Convolutional Neural Networks

Prabhanjans, P.; Punathil, A. N.; V K, A.; Thomas T, H. M.; Sasidharan, B. K.; Shaikh, H.; Varghese, A. J.; Kuchipudi, R. B.; Pavamani, S.; Rajan, J.

2026-03-13 radiology and imaging 10.64898/2026.03.12.26347996 medRxiv
Top 0.1%
26.6%
Show abstract

Head and neck cancer (HNC) requires accurate tumor delineation for effective radiotherapy planning. Manual segmentation of tumor regions is time-consuming and subject to considerable inter-observer variability. Although several automated approaches have been proposed, many rely on multimodal imaging such as PET/CT, which is expensive, less accessible in many clinical settings, and increases the burden on patients. In this work, we investigate a CT-only three-dimensional segmentation framework that provides a clinically practical and resource-efficient alternative. CT images of 136 head and neck cancer patients from the publicly available HN1 dataset in The Cancer Imaging Archive (TCIA) were used along with 30 additional cases from a private dataset collected at a tertiary care centre, Christian Medical College (CMC), Vellore, India. A fully automated segmentation model was developed to delineate the primary gross tumor volume (GTV) using the 3D nnU-Net framework. The models were trained using the HN1 dataset and an extended HN1+CMC dataset that included the additional private cases. Performance was evaluated using three-fold cross-validation with standard segmentation metrics including Dice Similarity Coefficient (DSC), Intersection over Union (IoU), and the 95th percentile Hausdorff Distance (HD95). The proposed CT-based model achieved a Global Dice of 0.63 and a Median Dice of 0.60 on the HN1 dataset. When the additional CMC cases were incorporated during training, the performance improved to a Global Dice of 0.65 and a Median Dice of 0.71. These results demonstrate that 3D nnU-Net can effectively segment head and neck tumors from CT images alone. The proposed CT-only approach provides a cost-effective and scalable solution that can support radiotherapy treatment planning and help reduce variability in clinical workflows.

11
Deep learning-assisted multiple organ segmentation from whole-body CT images

Salimi, y.; Shiri, I.; MAnsouri, Z.; Zaidi, H.

2023-10-21 radiology and imaging 10.1101/2023.10.20.23297331 medRxiv
Top 0.1%
26.3%
Show abstract

BackgroundAutomated organ segmentation from computed tomography (CT) images facilitates a number of clinical applications, including clinical diagnosis, monitoring of treatment response, quantification, radiation therapy treatment planning, and radiation dosimetry. PurposeTo develop a novel deep learning framework to generate multi-organ masks from CT images for 23 different body organs. MethodsA dataset consisting of 3106 CT images (649,398 axial 2D CT slices, 13,640 images/segment pairs) and ground-truth manual segmentation from various online available databases were collected. After cropping them to body contour, they were resized, normalized and used to train separate models for 23 organs. Data were split to train (80%) and test (20%) covering all the databases. A Res-UNET model was trained to generate segmentation masks from the input normalized CT images. The model output was converted back to the original dimensions and compared with ground-truth segmentation masks in terms of Dice and Jaccard coefficients. The information about organ positions was implemented during post-processing by providing six anchor organ segmentations as input. Our model was compared with the online available "TotalSegmentator" model through testing our model on their test datasets and their model on our test datasets. ResultsThe average Dice coefficient before and after post-processing was 84.28% and 83.26% respectively. The average Jaccard index was 76.17 and 70.60 before and after post-processing respectively. Dice coefficients over 90% were achieved for the liver, heart, bones, kidneys, spleen, femur heads, lungs, aorta, eyes, and brain segmentation masks. Post-processing improved the performance in only nine organs. Our model on the TotalSegmentator dataset was better than their models on our dataset in five organs out of 15 common organs and achieved almost similar performance for two organs. ConclusionsThe availability of a fast and reliable multi-organ segmentation tool leverages implementation in clinical setting. In this study, we developed deep learning models to segment multiple body organs and compared the performance of our models with different algorithms. Our model was trained on images presenting with large variability emanating from different databases producing acceptable results even in cases with unusual anatomies and pathologies, such as splenomegaly. We recommend using these algorithms for organs providing good performance. One of the main merits of our proposed models is their lightweight nature with an average inference time of 1.67 seconds per case per organ for a total-body CT image, which facilitates their implementation on standard computers.

12
Deep learning models to predict mammographic density jointly on standard dose and low dose images

Squires, S.; Mackenzie, A.; Evans, D. G.; Howell, S. J.; Astley, S. M.

2024-04-12 radiology and imaging 10.1101/2024.04.10.24305572 medRxiv
Top 0.1%
26.0%
Show abstract

ObjectivesMammographic density is associated with increased risk of developing breast cancer. Automated estimation of density in women below normal screening age would enable earlier risk stratification. We are piloting the use of low dose mammograms combined with models that can make accurate mammographic density estimates. MethodsThree models were trained on a joint set (107,619) of standard dose mammograms with associated density scores and their simulated low dose counterparts such that the models made predictions on standard and low dose mammograms. A second set of models was trained separately on the standard and simulated low dose mammograms. All models were tested on a held-out set from the training data and an independent dataset with 294 pairs of standard and real low dose mammograms. ResultsThe root mean squared errors (RMSE) between the model predictions and density scores on standard and simulated low dose images were 8.26 (8.16-8.36) and 8.27 (8.17-8.38) respectively. The RMSE between predictions on standard and simulated low dose images for the jointly trained models was 1.91 (1.88-1.96). The RMSE of the predictions on the real low dose images compared to the standard dose images is 3.79 (2.75-4.99). ConclusionsDeep learning models make density predictions on low dose images with similar quality as on standard dose images. Such automated analysis of low dose mammograms could contribute to accurate breast cancer risk estimation in younger women enabling stratification for further monitoring and preventative therapy. Advances in knowledgeMammographic density can be estimated in low dose mammograms with similar quality to standard dose mammograms.

13
FLASH Radiotherapy is faster than a heartbeat: A compartmental model to illustrate the interplay between tissue oxygen perfusion and ultra-high dose rate effects.

Ballesteros-Zebadua, P.; Jansen, J.; Grilij, V.; Franco-Perez, J.; Vozenin, M.-C.; Abolfath, R.

2026-03-16 biochemistry 10.64898/2026.03.12.711443 medRxiv
Top 0.1%
23.9%
Show abstract

Ultra-high-dose-rate therapy enhances the protection of normal tissues and reduces side effects while effectively controlling tumors. This biological phenomenon is called the FLASH effect, and when observed, therapy is called FLASH Radiotherapy (FLASH-RT). Various hypotheses have been proposed to explain how ultra-high dose rates achieve these effects under different conditions, with the impact of tissue oxygen perfusion still needing further investigation. FLASH-RT involves brief exposure to radiation, which results in fewer heartbeats occurring during the irradiation period, which could lead to reduced tissue oxygen perfusion occurring during the treatment timeframe. Therefore, we developed a compartmental model to simulate oxygen transfer and its interaction with radiation. The proposed model consists of three compartments: 1) the heart and arteries; 2) the irradiated brains blood vessels and capillaries; and 3) the irradiated brain tissue. We employed a system of differential equations, incorporating experimental data from in vivo oxygen measurements using the Oxyphor probe in the brain, to fit the model parameters to the experimental results. This model shows how dose rate and oxygen perfusion could influence chemical processes such as lipid peroxidation, potentially leading to differential biological effects. Our analysis of lipid peroxidation as a function of dose rate revealed a sigmoidal dose-rate-response curve that correlates well with several published biological response datasets. Our results indicate that the differential chemical effects of FLASH-RT compared with conventional dose rates may depend on factors such as oxygen perfusion, consumption, and tissue oxygen tension. This suggests that the temporal dynamics of oxygen could play a crucial role in enhancing the therapeutic window for FLASH-RT treatments. Furthermore, it suggests that the magnitude of some observed FLASH effects may vary across tissues or tumors and across experimental models, given differential oxygen dynamics.

14
Evaluating clinical acceptability of organ-at-risk segmentation In head & neck cancer using a compendium of open-source 3D convolutional neural networks

Marsilla, J.; Won Kim, J.; Kim, S.; Tkachuck, D.; Rey-McIntyre, K.; Patel, T.; Tadic, T.; Liu, F.-F.; Bratman, S.; Hope, A.; Haibe-Kains, B.

2022-01-25 radiology and imaging 10.1101/2022.01.15.22269276 medRxiv
Top 0.1%
22.8%
Show abstract

Background and PurposeAuto-segmentation of organs at risk (OAR) in cancer patients is essential for enhancing radiotherapy planning efficacy and reducing inter-observer variability. Deep learning auto-segmentation models have shown promise, but their lack of transparency and reproducibility hinders their generalizability and clinical acceptability, limiting their use in clinical settings. Materials and MethodsThis study introduces SCARF (auto-Segmentation Clinical Acceptability & Reproducibility Framework), a comprehensive six-stage reproducible framework designed to benchmark open-source convolutional neural networks for auto-segmentation of 19 essential OARs in head and neck cancer (HNC). ResultsSCARF offers an easily implementable framework for designing and reproducibly benchmarking auto-segmentation tools, along with thorough expert assessment capabilities. Expert assessment labelled 16/19 AI-generated OAR categories as acceptable with minor revisions. Boundary distance metrics, such as 95th Percentile Hausdorff Distance (95HD), were found to be 2x more correlated to Mean Acceptability Rating (MAR) than volumetric overlap metrics (DICE). ConclusionsThe introduction of SCARF, our auto-Segmentation Clinical Acceptability & Reproducibility Framework, represents a significant step forward in systematically assessing the performance of AI models for auto-segmentation in radiation therapy planning. By providing a comprehensive and reproducible framework, SCARF facilitates benchmarking and expert assessment of AI-driven auto-segmentation tools, addressing the need for transparency and reproducibility in this domain. The robust foundation laid by SCARF enables the progression towards the creation of usable AI tools in the field of radiation therapy. Through its emphasis on clinical acceptability and expert assessment, SCARF fosters the integration of AI models into clinical environments, paving the way for more randomised clinical trials to evaluate their real-world impact. O_TEXTBOXHighlightsO_LIOur study highlights the significance of both quantitative and qualitative controls for benchmarking new auto-segmentation systems effectively, promoting a more robust evaluation process of AI tools. C_LIO_LIWe address the lack of baseline models for medical image segmentation benchmarking by presenting SCARF, a comprehensive and reproducible six-stage framework, which serves as a valuable resource for advancing auto-segmentation research and contributing to the foundation of AI tools in radiation therapy planning. C_LIO_LISCARF enables benchmarking of 11 open-source convolutional neural networks (CNN) against 19 essential organs-at-risk (OARs) for radiation therapy in head and neck cancer, fostering transparency and facilitating external validation. C_LIO_LITo accurately assess the performance of auto-segmentation models, we introduce a clinical assessment toolkit based on the open-source QUANNOTATE platform, further promoting the use of external validation tools and expert assessment. C_LIO_LIOur study emphasises the importance of clinical acceptability testing and advocates its integration into developing validated AI tools for radiation therapy planning and beyond, bridging the gap between AI research and clinical practice. C_LI C_TEXTBOX

15
Evaluating the Large Language Model-Based Quality Assurance Tool for Auto-Contouring

Tozuka, R.; Akita, T.; Matsuda, M.; Tanno, H.; Saito, M.; Nemoto, H.; Mitsuda, K.; Kadoya, N.; Jingu, K.; Onishi, H.

2026-04-01 radiology and imaging 10.64898/2026.03.31.26349802 medRxiv
Top 0.1%
22.3%
Show abstract

Purpose: Manual verification of AI-based auto-contouring is labor-intensive and prone to fatigue-related errors. This study developed the large language model (LLM)-based automated Quality Assurance (QA) for auto-contouring (LAQUA) system using a multimodal LLM, Gemini 2.5 Pro, and evaluated its feasibility as a clinical primary screening tool to streamline the QA workflow. Methods: Twenty male pelvic CT scans from an open dataset were utilized. Three distinct auto-contouring software packages (OncoStudio, RatoGuide prototype and syngo.via) were evaluated. Auto-contouring results for each slice were exported as PDF images with overlaid contours and input into Gemini 2.5 Pro. The LLM was instructed to rate the contour quality on a 5-point clinical scale (5: Optimal; 4: Acceptable; 3: Suboptimal; 2: Unacceptable; redraw from scratch; 1: Unacceptable; organ not detected). Using evaluations by two board-certified radiation oncologists as ground truth, Spearman's rank correlation coefficients ({rho}) and weighted kappa coefficients ({kappa}) were calculated. Additionally, to assess screening performance, sensitivity and specificity were calculated by dichotomizing the scores into "Pass" and "Fail" using two different cutoffs (scores [&ge;] 3 and [&ge;] 4 as "Pass"). Finally, the alignment of the rationales provided by the LLM with the auto-contouring quality was evaluated by two board-certified radiation oncologists. This was conducted using a Likert scale assessing four domains (error detection, hallucination, clinical relevance, and anatomical understanding), each scored out of 2 points. Results: The LAQUA system demonstrated moderate to strong agreement with expert judgments across all evaluated organs ({rho}: 0.567 - 0.835; quadratic weighted {kappa} : 0.639 - 0.804), with the rectum showing the highest correlation. Regarding screening performance, a cutoff of [&ge;]3 as "Pass" achieved the highest sensitivity and specificity in specific subgroups, but with wide 95% confidence intervals (CIs). A cutoff of [&ge;]4 as "Pass" narrowed the CIs, yielding the highest sensitivity in the rectum (0.976) and the highest specificity in the left femoral head (0.933). Qualitatively, the LLM's rationales achieved an overall mean score of 1.70 {+/-} 0.48 (out of 2), with 155 of 291 outputs receiving perfect scores across all criteria. Conclusions: The LAQUA system demonstrated substantial agreement with expert evaluations in AI-based auto-contouring quality assessment. While potential overestimation bias (risk of missing "Fail" cases) warrants caution, the observed sensitivity suggests its feasibility as a primary screening QA tool to efficiently filter acceptable contours, thereby reducing the clinical workload.

16
PixelPrint: Three-dimensional printing of realistic patient-specific lung phantoms for validation of computed tomography post-processing and inference algorithms

Shapira, N.; Donovan, K.; Mei, K.; Geagan, M.; Roshkovan, L.; Gang, G.; Abed, M.; Linna, N. B.; Cranston, C. P.; Leary, C. N.; Dhanaliwala, A. H.; Kontos, D.; Litt, H. I.; Stayman, J. W.; Shinohara, R. T.; Noël, P. B.

2022-05-10 radiology and imaging 10.1101/2022.05.06.22274739 medRxiv
Top 0.1%
21.8%
Show abstract

BackgroundRadiomics and other modern clinical decision-support algorithms are emerging as the next frontier for diagnostic and prognostic medical imaging. However, heterogeneities in image characteristics due to variations in imaging systems and protocols hamper the advancement of reproducible feature extraction pipelines. There is a growing need for realistic patient-based phantoms that accurately mimic human anatomy and disease manifestations to provide consistent ground-truth targets when comparing different feature extraction or image cohort normalization techniques. Materials and MethodsPixelPrint was developed for 3D-printing lifelike lung phantoms for computed tomography (CT) by directly translating clinical images into printer instructions that control the density on a voxel-by-voxel basis. CT datasets of three COVID-19 pneumonia patients served as input for 3D-printing lung phantoms. Five radiologists rated patient and phantom images for imaging characteristics and diagnostic confidence in a blinded reader study. Linear mixed models were utilized to evaluate effect sizes of evaluating phantom as opposed to patient images. Finally, PixelPrints reproducibility was evaluated by producing four phantoms from the same clinical images. ResultsEstimated mean differences between patient and phantom images were small (0.03-0.29, using a 1-5 scale). Effect size assessment with respect to rating variabilities revealed that the effect of having a phantom in the image is within one-third of the inter- and intra-reader variabilities. PixelPrints production reproducibility tests showed high correspondence among four phantoms produced using the same patient images, with higher similarity scores between high-dose scans of the different phantoms than those measured between clinical-dose scans of a single phantom. ConclusionsWe demonstrated PixelPrints ability to produce lifelike 3D-printed CT lung phantoms reliably. These can provide ground-truth targets for validating the generalizability of inference-based decision-support algorithms between different health centers and imaging protocols, as well as for optimizing scan protocols with realistic patient-based phantoms.

17
Head and Neck Cancer Primary Tumor Auto Segmentation using Model Ensembling of Deep Learning in PET-CT Images

Naser, M. A.; Wahid, K. A.; van Dijk, L. V.; He, R.; Abdelaal, M. A.; Dede, C.; Mohamed, A. S. R.; Fuller, C. D.

2021-10-18 radiology and imaging 10.1101/2021.10.14.21264953 medRxiv
Top 0.1%
21.7%
Show abstract

Auto-segmentation of primary tumors in oropharyngeal cancer using PET/CT images is an unmet need that has the potential to improve radiation oncology workflows. In this study, we develop a series of deep learning models based on a 3D Residual Unet (ResUnet) architecture that can segment oropharyngeal tumors with high performance as demonstrated through internal and external validation of large-scale datasets (training size = 224 patients, testing size = 101 patients) as part of the 2021 HECKTOR Challenge. Specifically, we leverage ResUNet models with either 256 or 512 bottleneck layer channels that are able to demonstrate internal validation (10-fold cross-validation) mean Dice similarity coefficient (DSC) up to 0.771 and median 95% Hausdorff distance (95% HD) as low as 2.919 mm. We employ label fusion ensemble approaches, including Simultaneous Truth and Performance Level Estimation (STAPLE) and a voxel-level threshold approach based on majority voting (AVERAGE), to generate consensus segmentations on the test data by combining the segmentations produced through different trained cross-validation models. We demonstrate that our best performing ensembling approach (256 channels AVERAGE) achieves a mean DSC of 0.770 and median 95% HD of 3.143 mm through independent external validation on the test set. Concordance of internal and external validation results suggests our models are robust and can generalize well to unseen PET/CT data. We advocate that ResUNet models coupled to label fusion ensembling approaches are promising candidates for PET/CT oropharyngeal primary tumors auto-segmentation, with future investigations targeting the ideal combination of channel combinations and label fusion strategies to maximize segmentation performance.

18
Deep Learning-powered CT-less Multi-tracer Organ Segmentation from PET Images: A solution for unreliable CT segmentation in PET/CT Imaging

Salimi, Y.; Mansouri, Z.; Shiri, I.; Mainta, I.; Zaidi, H.

2024-08-28 radiology and imaging 10.1101/2024.08.27.24312482 medRxiv
Top 0.1%
19.6%
Show abstract

IntroductionThe common approach for organ segmentation in hybrid imaging relies on co-registered CT (CTAC) images. This method, however, presents several limitations in real clinical workflows where mismatch between PET and CT images are very common. Moreover, low-dose CTAC images have poor quality, thus challenging the segmentation task. Recent advances in CT-less PET imaging further highlight the necessity for an effective PET organ segmentation pipeline that does not rely on CT images. Therefore, the goal of this study was to develop a CT-less multi-tracer PET segmentation framework. MethodsWe collected 2062 PET/CT images from multiple scanners. The patients were injected with either 18F-FDG (1487) or 68Ga-PSMA (575). PET/CT images with any kind of mismatch between PET and CT images were detected through visual assessment and excluded from our study. Multiple organs were delineated on CT components using previously trained in-house developed nnU-Net models. The segmentation masks were resampled to co-registered PET images and used to train four different deep-learning models using different images as input, including non-corrected PET (PET-NC) and attenuation and scatter-corrected PET (PET-ASC) for 18F-FDG (tasks #1 and #2, respectively using 22 organs) and PET-NC and PET-ASC for 68Ga tracers (tasks #3 and #4, respectively, using 15 organs). The models performance was evaluated in terms of Dice coefficient, Jaccard index, and segment volume difference. ResultsThe average Dice coefficient over all organs was 0.81{+/-}0.15, 0.82{+/-}0.14, 0.77{+/-}0.17, and 0.79{+/-}0.16 for tasks #1, #2, #3, and #4, respectively. PET-ASC models outperformed PET-NC models (P-value < 0.05). The highest Dice values were achieved for the brain (0.93 to 0.96 in all four tasks), whereas the lowest values were achieved for small organs, such as the adrenal glands. The trained models showed robust performance on dynamic noisy images as well. ConclusionDeep learning models allow high performance multi-organ segmentation for two popular PET tracers without the use of CT information. These models may tackle the limitations of using CT segmentation in PET/CT image quantification, kinetic modeling, radiomics analysis, dosimetry, or any other tasks that require organ segmentation masks.

19
Dual-energy computed tomography imaging with megavoltage and kilovoltage x-ray spectra

Jadick, G.; Schlafly, G.; La Riviere, P.

2023-06-29 radiology and imaging 10.1101/2023.06.22.23291766 medRxiv
Top 0.1%
19.6%
Show abstract

PurposeSingle-energy computed tomography (CT) often suffers from poor contrast, yet it remains critical for effec-tive radiotherapy treatment. Modern therapy systems are often equipped with both megavoltage (MV) and kilovoltage (kV) x-ray sources and thus already possess the hardware needed for dual-energy (DE) CT. There exists an unexplored potential for enhanced image contrast using MV-kV DE-CT in radiotherapy contexts. ApproachA toy model comprising a single-line integral through a two-material object was designed for computing basis material signal-to-noise ratio (SNR) using estimation theory. Five dose-matched spectra (three kV, two MV) and three variables were considered: spectral combination, spectral dose allocation, and object material composition. The single-line model was extended to a simulated fan-beam CT acquisition of an anthropomorphic phantom with and without a metal implant. Basis material sinograms were computed and synthesized into virtual monoenergetic images (VMIs). MV-kV and kV-kV VMIs were compared with single-energy images. ResultsThe 80kV-140kV pair typically yielded the best SNRs, but for bone thicknesses greater than 8 cm, the detunedMV-80kV pair surpassed it. Peak MV-kV SNR was achieved with approximately 90% dose allocated to the MV spectrum. For the CT simulations, MV-kV VMIs yielded a higher contrast-to-noise ratio (CNR) than single-energy CT at specific monoenergies. With the metal implant, MV-kV produced a higher maximum CNR and lower minimum root-mean-square-error than kV-kV. ConclusionsThis work quantitatively analyzes MV-kV DE-CT imaging and assesses its potential advantages. This technique may yield improved contrast and accuracy relative to dose-matched single-energy CT or kV-kV DE-CT, depending on object composition.

20
A Real-World Evaluation of Failure Detection for Liver CT Segmentation

Bennett, J.; Woodland, M.; Castelo, A.; Altaie, M.; Antony, A.; Siddiqi, N. S.; Long, J. P.; Brock, K. K.

2026-06-29 radiology and imaging 10.64898/2026.06.26.26356692 medRxiv
Top 0.1%
19.3%
Show abstract

Deep learning models deployed in clinical imaging frequently encounter distribution shifts, yet most out-of-distribution (OOD) detection methods are evaluated only on controlled research datasets. As a result, it is unclear whether existing approaches can reliably identify segmentation failures that arise in real-world clinical practice. We evaluated six OOD detection methods on a deployed liver CT segmentation model (3D nnU-Net) using internal data from 400 patients and external data from 100 patients collected across nearly 70 sites in 7 countries. One method was Pairwise Surface DSC, a surface-based extension of Pairwise DSC, that we introduced. OOD performance was measured using sensitivity, AUROC, and balanced accuracy, with thresholds determined on an independent cohort of 400 patients using the Youden J statistic. Statistical significance was assessed using McNemar tests and stratified bootstraps ( = 0.05) with Benjamini-Hochberg correction. Pairwise Surface DSC was the top-performing method, with perfect sensitivities (1.00), near-perfect AUROCs (0.97 internal; 1.00 external), and the highest balanced accuracies (0.94 internal; 0.88 external; p<0.001). These results show that automated failure detection for liver CT segmentation is clinically feasible and that Pairwise Surface DSC is a promising candidate for deployment. Our code is available at https://github.com/mckellwoodland/liver_ct_ood_translation.