Quality versus quantity of training datasets for artificial intelligence-based whole liver segmentation
Castelo, A.; O'Connor, C.; Gupta, A. C.; Anderson, B. M.; Woodland, M.; Altaie, M.; Koay, E. J.; Odisio, B. C.; Tang, T. T.; Brock, K. K.
Show abstract
Artificial intelligence (AI) based segmentation has many medical applications but limited curated datasets challenge model training; this study compares the impact of dataset annotation quality and quantity on whole liver AI segmentation performance. We obtained 3,089 abdominal computed tomography scans with whole-liver contours from MD Anderson Cancer Center (MDA) and a MICCAI challenge. A total of 249 scans were withheld for testing of which 30, MICCAI challenge data, were reserved for external validation. The remaining scans were divided into mixed-curation and highly-curated groups, randomly sampled into sub-datasets of various sizes, and used to train 3D nnU-Net segmentation models. Dice similarity coefficients (DSC), surface DSC with 2mm margins (SD 2mm), the 95th percentile of Hausdorff distance (HD95), and 2D axial slice DSC (Slice DSC) were used to evaluate model performance. The highly curated, 244-scan model (DSC=0.971, SD 2mm=0.958, HD95=2.98mm) performed insignificantly different on 3D evaluation metrics to the mixed-curation 2,840-scan model (DSC=0.971 [p>.999], SD 2mm=0.958 [p>.999], HD95=2.87mm [p>.999]). The 710-scan mixed-curation (Slice DSC=0.929) significantly outperformed the highly curated, 244-scan model (Slice DSC=0.923 [p=0.012]) on the 30 external scans. Highly curated datasets yielded equivalent performance to datasets that were a full order of magnitude larger. The benefits of larger, mixed-curation datasets are evidenced in model generalizability metrics and local improvements. In conclusion, tradeoffs between dataset quality and quantity for model training are nuanced and goal dependent.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Phase Recognition in Contrast-Enhanced CT Scans based on Deep Learning and Random Sampling 96%
- Necessity and Impact of Specialization of Large Foundation Model for Medical Segmentation Tasks 95%
- Fully Automated Explainable Abdominal CT Contrast Media Phase Classification Using Organ Segmentation and Machine Learning 94%
Similar papers in this journal
- Quantifying Epistemic and Aleatoric Uncertainty in 3D U-Net Segmentation 93%
- The Effect of Image Resolution on Automated Classification of Chest X-rays 93%
- “E Pluribus Unum”: Prospective acceptability benchmarking from the Contouring Collaborative for Consensus in Radiation Oncology (C3RO) Crowdsourced Initiative for Multi-Observer Segmentation 93%
Similar papers in this journal
- Evaluating Large Language Model-Generated Brain MRI Protocols: Performance of GPT4o, o3-mini, DeepSeek-R1 and Qwen2.5-72B 94%
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 93%
- Fully automatic segmentation of craniomaxillofacial CT scans for computer-assisted orthognathic surgery planning using the nnU-Net framework 93%
Similar papers in this journal
- Effective Deep Learning Approaches for Predicting COVID-19 Outcomes from Chest Computed Tomography Volumes 94%
- Improving Rectal Tumor Segmentation with Anomaly Fusion Derived from Anatomical Inpainting: A Multicenter Study 94%
- Inconsistency of AI in Intracranial Aneurysm Detection with Varying Dose and Image Reconstruction 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.