Label-Free Threshold Selection for Out-of-Distribution Detection in Liver CT Segmentation
Nielsen, M.; Castelo, A.; Altaie, M.; Bennett, J.; Anthony, A.; Siddiqi, N. S.; Gupta, A. C.; Brock, K. K.; Woodland, M.
Show abstract
Reliable clinical deployment of automated liver segmentation requires mechanisms for detecting failures in rare and previously unseen scenarios. Achieving this goal requires an appropriately calibrated threshold that converts an out-of-distribution (OOD) score into a failure prediction. However, threshold calibration typically relies on expert-labeled failures, creating a substantial annotation burden when failures are rare. Building upon our prior work, which uses Pairwise Surface DSC scores as indicators of segmentation quality, we propose a label-free framework for calibrating OOD score thresholds. First, we fitted a log-t distribution to Pairwise Surface DSC scores from a validation set of 400 internal scans to approximate an in-distribution score distribution. New segmentations were assigned significance scores based on their extremity under this fitted distribution and categorized into Low, Medium, and High Risk review groups using statistically principled cutoffs of 0.25 and 0.05. The fitted log-t distribution provided a strong fit to the observed scores and remained robust to moderate contamination by OOD cases. On an independent test set of 500 internal and external scans, the combined Medium and High Risk categories achieved 100% sensitivity and 79% specificity, whereas the High Risk category alone achieved 78% sensitivity and 96% specificity. These results indicate that clinically meaningful failure detection can be derived from unlabeled data. Our code is available at https://github.com/marshalln7/Label_Free_OOD_Threshold_Selection.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ENRICHing Medical Imaging Training Sets Enables More Efficient Machine Learning 93%
- Quantification of abdominal fat from computed tomography using deep learning and its association with electronic health records in an academic biobank 93%
- Analysis of Eligibility Criteria Clusters Based on Large Language Models for Clinical Trial Design 92%
Similar papers in this journal
- Weakly supervised classification of rare aortic valve malformations using unlabeled cardiac MRI sequences 92%
- An Ensembled Deep Learning Model Outperforms Human Experts in Diagnosing Biliary Atresia from Sonographic Gallbladder Images 92%
- Artificial Intelligence System Reduces False-Positive Findings in the Interpretation of Breast Ultrasound Exams 91%
Similar papers in this journal
Similar papers in this journal
- Developing a Fully Automated Imaging Biomarker for HCC Risk Assessment via MRI-Based Tumor Segmentation and EPM 92%
- Generating synthetic data in digital pathology through diffusion models: a multifaceted approach to evaluation 92%
- On evaluation metrics for medical applications of artificial intelligence 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.