Back

AUGMENT: a framework for robust assessment of the clinical utility of segmentation algorithms

McCague, C.; Buddenkotte, T.; Escudero Sanchez, L.; Hulse, D.; Pintican, R.; Rundo, L.; Freeman, S.; Nougaret, S.; Rizzo, S.; Loughborough, W.; Andreou, A.; Parsons, C.; Piyatissa, P.; Aloysius, T.; Mouritsen Luxhoj, C.; Aniq, I.; James, S.; Dhesi, B.; De Paepe, K.; Tanner, J.; Abulaban, O.; Lee, J.; Majcher, V.; O Sullivan, M.; Celli, V.; Colarieti, A.; Samoshkin, A.; Carcani, E.; Ramlee, S.; Al Sad, M. S.; Doran, S. J.; Cho, W.; DArcy, J.; Brenton, J. D.; Couturier, D. L.; Oktem, O.; Woitek, R.; Schoenlieb, C. B.; Sala, E.; Crispin Ortuzar, M.

2024-09-23 radiology and imaging
10.1101/2024.09.20.24313970 medRxiv
Show abstract

BackgroundEvaluating AI-based segmentation models primarily relies on quantitative metrics, but it remains unclear if this approach leads to practical, clinically applicable tools. PurposeTo create a systematic framework for evaluating the performance of segmentation models using clinically relevant criteria. Materials and MethodsWe developed the AUGMENT framework (Assessing Utility of seGMENtation Tools), based on a structured classification of main categories of error in segmentation tasks. To evaluate the framework, we assembled a team of 20 clinicians covering a broad range of radiological expertise and analysed the challenging task of segmenting metastatic ovarian cancer using AI. We used three evaluation methods: (i) Dice Similarity Coefficient (DSC), (ii) visual Turing test, assessing 429 segmented disease-sites on 80 CT scans from the Cancer Imaging Atlas), and (iii) AUGMENT framework, where 3 radiologists and the AI-model created segmentations of 784 separate disease sites on 27 CT scans from a multi-institution dataset. ResultsThe AI model had modest technical performance (DSC=72{+/-}19 for the pelvic and ovarian disease, and 64{+/-}24 for omental disease), and it failed the visual Turing test. However, the AUGMENT framework revealed that (i) the AI model produced segmentations of the same quality as radiologists (p=.46), and (ii) it enabled radiologists to produce human+AI collaborative segmentations of significantly higher quality (p=<.001) and in significantly less time (p=<.001). ConclusionQuantitative performance metrics of segmentation algorithms can mask their clinical utility. The AUGMENT framework enables the systematic identification of clinically usable AI-models and highlights the importance of assessing the interaction between AI tools and radiologists. Summary statementOur framework, called AUGMENT, provides an objective assessment of the clinical utility of segmentation algorithms based on well-established error categories. Key resultsO_LICombining quantitative metrics with qualitative information on performance from domain experts whose work is impacted by an algorithms use is a more accurate, transparent and trustworthy way of appraising an algorithm than using quantitative metrics alone. C_LIO_LIThe AUGMENT framework captures clinical utility in terms of segmentation quality and human+AI complementarity even in algorithms with modest technical segmentation performance. C_LIO_LIAUGMENT might have utility during the development and validation process, including in segmentation challenges, for those seeking clinical translation, and to audit model performance after integration into clinical practice. C_LI

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.