Evaluating the Large Language Model-Based Quality Assurance Tool for Auto-Contouring
Tozuka, R.; Akita, T.; Matsuda, M.; Tanno, H.; Saito, M.; Nemoto, H.; Mitsuda, K.; Kadoya, N.; Jingu, K.; Onishi, H.
Show abstract
Purpose: Manual verification of AI-based auto-contouring is labor-intensive and prone to fatigue-related errors. This study developed the large language model (LLM)-based automated Quality Assurance (QA) for auto-contouring (LAQUA) system using a multimodal LLM, Gemini 2.5 Pro, and evaluated its feasibility as a clinical primary screening tool to streamline the QA workflow. Methods: Twenty male pelvic CT scans from an open dataset were utilized. Three distinct auto-contouring software packages (OncoStudio, RatoGuide prototype and syngo.via) were evaluated. Auto-contouring results for each slice were exported as PDF images with overlaid contours and input into Gemini 2.5 Pro. The LLM was instructed to rate the contour quality on a 5-point clinical scale (5: Optimal; 4: Acceptable; 3: Suboptimal; 2: Unacceptable; redraw from scratch; 1: Unacceptable; organ not detected). Using evaluations by two board-certified radiation oncologists as ground truth, Spearman's rank correlation coefficients ({rho}) and weighted kappa coefficients ({kappa}) were calculated. Additionally, to assess screening performance, sensitivity and specificity were calculated by dichotomizing the scores into "Pass" and "Fail" using two different cutoffs (scores [≥] 3 and [≥] 4 as "Pass"). Finally, the alignment of the rationales provided by the LLM with the auto-contouring quality was evaluated by two board-certified radiation oncologists. This was conducted using a Likert scale assessing four domains (error detection, hallucination, clinical relevance, and anatomical understanding), each scored out of 2 points. Results: The LAQUA system demonstrated moderate to strong agreement with expert judgments across all evaluated organs ({rho}: 0.567 - 0.835; quadratic weighted {kappa} : 0.639 - 0.804), with the rectum showing the highest correlation. Regarding screening performance, a cutoff of [≥]3 as "Pass" achieved the highest sensitivity and specificity in specific subgroups, but with wide 95% confidence intervals (CIs). A cutoff of [≥]4 as "Pass" narrowed the CIs, yielding the highest sensitivity in the rectum (0.976) and the highest specificity in the left femoral head (0.933). Qualitatively, the LLM's rationales achieved an overall mean score of 1.70 {+/-} 0.48 (out of 2), with 155 of 291 outputs receiving perfect scores across all criteria. Conclusions: The LAQUA system demonstrated substantial agreement with expert evaluations in AI-based auto-contouring quality assessment. While potential overestimation bias (risk of missing "Fail" cases) warrants caution, the observed sensitivity suggests its feasibility as a primary screening QA tool to efficiently filter acceptable contours, thereby reducing the clinical workload.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Quality Assurance Assessment of Intra-Acquisition Diffusion-Weighted and T2-Weighted Magnetic Resonance Imaging Registration and Contour Propagation for Head and Neck Cancer Radiotherapy 95%
- Cell lines of the same anatomic site and histologic type show large variability in intrinsic radiosensitivity and relative biological effectiveness to protons and carbon ions 94%
- Evaluation of inverse treatment planning for Gamma Knife radiosurgery using fMRI brain activation maps as organs at risk 93%
Similar papers in this journal
- Comprehensive Quantitative Evaluation of Inter-observer Delineation Performance of MR-guided Delineation of Oropharyngeal Gross Tumor Volumes and High-risk Clinical Target Therapy: An R-IDEAL Stage 0 Prospective Study 97%
- Cluster-Based Toxicity Estimation of Osteoradionecrosis via Unsupervised Machine Learning: Moving Beyond Single Dose-Parameter Normal Tissue Complication Probability by Using Whole Dose-Volume Histograms for Cohort Risk Stratification 96%
- Clinical Impact of Contouring Variability for Prostate Cancer Tumor Boost 96%
Similar papers in this journal
Similar papers in this journal
- Auto-Detection and Segmentation of Involved Lymph Nodes in HPV-Associated Oropharyngeal Cancer Using a Convolutional Deep Learning Neural Network 95%
- Development of a High-Performance Multiparametric MRI Oropharyngeal Primary Tumor Auto-Segmentation Deep Learning Model and Investigation of Input Channel Effects: Results from a Prospective Imaging Registry 94%
- Morphological changes after cranial fractionated photon radiotherapy: localized loss of white matter and grey matter volume with increasing dose 94%
Similar papers in this journal
- Artificial Intelligence Uncertainty Quantification in Radiotherapy Applications - A Scoping Review 96%
- Evaluation of indirect damage and damage saturation effects in dose-response curves of hypofractionated radiotherapy of early-stage NSCLC and brain metastases 94%
- Deep learning NTCP model for late dysphagia after radiotherapy for head and neck cancer patients based on 3D dose, CT and segmentations 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.