Evaluating clinical acceptability of organ-at-risk segmentation In head & neck cancer using a compendium of open-source 3D convolutional neural networks
Marsilla, J.; Won Kim, J.; Kim, S.; Tkachuck, D.; Rey-McIntyre, K.; Patel, T.; Tadic, T.; Liu, F.-F.; Bratman, S.; Hope, A.; Haibe-Kains, B.
Show abstract
Background and PurposeAuto-segmentation of organs at risk (OAR) in cancer patients is essential for enhancing radiotherapy planning efficacy and reducing inter-observer variability. Deep learning auto-segmentation models have shown promise, but their lack of transparency and reproducibility hinders their generalizability and clinical acceptability, limiting their use in clinical settings. Materials and MethodsThis study introduces SCARF (auto-Segmentation Clinical Acceptability & Reproducibility Framework), a comprehensive six-stage reproducible framework designed to benchmark open-source convolutional neural networks for auto-segmentation of 19 essential OARs in head and neck cancer (HNC). ResultsSCARF offers an easily implementable framework for designing and reproducibly benchmarking auto-segmentation tools, along with thorough expert assessment capabilities. Expert assessment labelled 16/19 AI-generated OAR categories as acceptable with minor revisions. Boundary distance metrics, such as 95th Percentile Hausdorff Distance (95HD), were found to be 2x more correlated to Mean Acceptability Rating (MAR) than volumetric overlap metrics (DICE). ConclusionsThe introduction of SCARF, our auto-Segmentation Clinical Acceptability & Reproducibility Framework, represents a significant step forward in systematically assessing the performance of AI models for auto-segmentation in radiation therapy planning. By providing a comprehensive and reproducible framework, SCARF facilitates benchmarking and expert assessment of AI-driven auto-segmentation tools, addressing the need for transparency and reproducibility in this domain. The robust foundation laid by SCARF enables the progression towards the creation of usable AI tools in the field of radiation therapy. Through its emphasis on clinical acceptability and expert assessment, SCARF fosters the integration of AI models into clinical environments, paving the way for more randomised clinical trials to evaluate their real-world impact. O_TEXTBOXHighlightsO_LIOur study highlights the significance of both quantitative and qualitative controls for benchmarking new auto-segmentation systems effectively, promoting a more robust evaluation process of AI tools. C_LIO_LIWe address the lack of baseline models for medical image segmentation benchmarking by presenting SCARF, a comprehensive and reproducible six-stage framework, which serves as a valuable resource for advancing auto-segmentation research and contributing to the foundation of AI tools in radiation therapy planning. C_LIO_LISCARF enables benchmarking of 11 open-source convolutional neural networks (CNN) against 19 essential organs-at-risk (OARs) for radiation therapy in head and neck cancer, fostering transparency and facilitating external validation. C_LIO_LITo accurately assess the performance of auto-segmentation models, we introduce a clinical assessment toolkit based on the open-source QUANNOTATE platform, further promoting the use of external validation tools and expert assessment. C_LIO_LIOur study emphasises the importance of clinical acceptability testing and advocates its integration into developing validated AI tools for radiation therapy planning and beyond, bridging the gap between AI research and clinical practice. C_LI C_TEXTBOX
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Quality Assurance Assessment of Intra-Acquisition Diffusion-Weighted and T2-Weighted Magnetic Resonance Imaging Registration and Contour Propagation for Head and Neck Cancer Radiotherapy 97%
- PSMA-Hornet: fully-automated, multi-target segmentation of healthy organs in PSMA PET/CT images 94%
- Persistent Homology of Tumor CT Scans is Associated with Survival In Lung Cancer 94%
Similar papers in this journal
- Identifying relationships between imaging phenotypes and lung cancer-related mutation status: EGFR and KRAS 94%
- Intermittent radiotherapy as alternative treatment for recurrent high grade glioma: A modelling study based on longitudinal tumor measurements 94%
- Dual Adversarial Deconfounding Autoencoder for joint batch-effects removal from multi-center and multi-scanner radiomics data 93%
Similar papers in this journal
- Large language models to help appeal denied radiotherapy services 94%
- Histology-based Prediction of Therapy Response to Neoadjuvant Chemotherapy for Esophageal and Esophagogastric Junction Adenocarcinomas Using Deep Learning 93%
- Using Adversarial Images to Assess the Stability of Deep Learning Models Trained on Diagnostic Images in Oncology 90%
Similar papers in this journal
Similar papers in this journal
- Large-scale crowdsourced radiotherapy segmentations across a variety of cancer anatomic sites: Interobserver expert/non-expert and multi-observer composite tumor and normal tissue delineation annotations from a prospective educational challenge 97%
- Segmentation of vestibular schwannoma from MRI — An open annotated dataset and baseline algorithm 97%
- Weekly Intra-Treatment Diffusion Weighted Imaging Dataset for Head and Neck Cancer Patients Undergoing MR-linac Treatment 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.