AUGMENT: a framework for robust assessment of the clinical utility of segmentation algorithms
McCague, C.; Buddenkotte, T.; Escudero Sanchez, L.; Hulse, D.; Pintican, R.; Rundo, L.; Freeman, S.; Nougaret, S.; Rizzo, S.; Loughborough, W.; Andreou, A.; Parsons, C.; Piyatissa, P.; Aloysius, T.; Mouritsen Luxhoj, C.; Aniq, I.; James, S.; Dhesi, B.; De Paepe, K.; Tanner, J.; Abulaban, O.; Lee, J.; Majcher, V.; O Sullivan, M.; Celli, V.; Colarieti, A.; Samoshkin, A.; Carcani, E.; Ramlee, S.; Al Sad, M. S.; Doran, S. J.; Cho, W.; DArcy, J.; Brenton, J. D.; Couturier, D. L.; Oktem, O.; Woitek, R.; Schoenlieb, C. B.; Sala, E.; Crispin Ortuzar, M.
Show abstract
BackgroundEvaluating AI-based segmentation models primarily relies on quantitative metrics, but it remains unclear if this approach leads to practical, clinically applicable tools. PurposeTo create a systematic framework for evaluating the performance of segmentation models using clinically relevant criteria. Materials and MethodsWe developed the AUGMENT framework (Assessing Utility of seGMENtation Tools), based on a structured classification of main categories of error in segmentation tasks. To evaluate the framework, we assembled a team of 20 clinicians covering a broad range of radiological expertise and analysed the challenging task of segmenting metastatic ovarian cancer using AI. We used three evaluation methods: (i) Dice Similarity Coefficient (DSC), (ii) visual Turing test, assessing 429 segmented disease-sites on 80 CT scans from the Cancer Imaging Atlas), and (iii) AUGMENT framework, where 3 radiologists and the AI-model created segmentations of 784 separate disease sites on 27 CT scans from a multi-institution dataset. ResultsThe AI model had modest technical performance (DSC=72{+/-}19 for the pelvic and ovarian disease, and 64{+/-}24 for omental disease), and it failed the visual Turing test. However, the AUGMENT framework revealed that (i) the AI model produced segmentations of the same quality as radiologists (p=.46), and (ii) it enabled radiologists to produce human+AI collaborative segmentations of significantly higher quality (p=<.001) and in significantly less time (p=<.001). ConclusionQuantitative performance metrics of segmentation algorithms can mask their clinical utility. The AUGMENT framework enables the systematic identification of clinically usable AI-models and highlights the importance of assessing the interaction between AI tools and radiologists. Summary statementOur framework, called AUGMENT, provides an objective assessment of the clinical utility of segmentation algorithms based on well-established error categories. Key resultsO_LICombining quantitative metrics with qualitative information on performance from domain experts whose work is impacted by an algorithms use is a more accurate, transparent and trustworthy way of appraising an algorithm than using quantitative metrics alone. C_LIO_LIThe AUGMENT framework captures clinical utility in terms of segmentation quality and human+AI complementarity even in algorithms with modest technical segmentation performance. C_LIO_LIAUGMENT might have utility during the development and validation process, including in segmentation challenges, for those seeking clinical translation, and to audit model performance after integration into clinical practice. C_LI
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 95%
- Evaluating Large Language Model-Generated Brain MRI Protocols: Performance of GPT4o, o3-mini, DeepSeek-R1 and Qwen2.5-72B 94%
- Impact of Non-Contrast Enhanced Imaging Input Sequences on the Generation of Virtual Contrast-Enhanced Breast MRI Scans using Neural Networks 93%
Similar papers in this journal
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 94%
- Development and Validation of a Deep Learning Model for Detecting Signs of Tuberculosis on Chest Radiographs among US-bound Immigrants and Refugees 94%
- Assessing generalizability of an AI-based visual test for cervical cancer screening 94%
Similar papers in this journal
- Classification performance bias between training and test sets in a limited mammography dataset 95%
- Enhancing Semantic Segmentation in Chest X-Ray Images through Image Preprocessing: ps-KDE for Pixel-wise Substitution by Kernel Density Estimation 94%
- ai-corona : Radiologist-Assistant Deep Learning Framework for COVID-19 Diagnosis in Chest CT Scans 92%
Similar papers in this journal
- Content-based image retrieval assists radiologists in diagnosing eye and orbital mass lesions in MRI 95%
- Reproducible And Clinically Translatable Deep Neural Networks For Cervical Screening 94%
- Generating synthetic data in digital pathology through diffusion models: a multifaceted approach to evaluation 93%
Similar papers in this journal
- A human-in-the-loop explanation framework for morphologically transparent AI predictions from whole-slide images 93%
- Public evidence on AI products for digital pathology 92%
- Diagnostic Accuracy of Artificial Intelligence in Classifying HER2 Status in Breast Cancer Immunohistochemistry Slides and Implications for HER2-Low Cases: A Systematic Review and Meta-Analysis 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.