Back

Human In the Loop Challenges for Quality Annotation of Pre-Cancer Lesions in Clinical Oral Images

Mandal, S.; Mendonca, P.; Gurushanth, K.; Thakur, H.; Birur, P.; Shetty, A.; Pal, D.

2026-07-04 dentistry and oral medicine
10.64898/2026.07.02.26355859 medRxiv
Show abstract

Background: The hyperplasia and dysplasia stage (pre-cancer) offers a viable opportunity to reduce the incidence and mortality of oral cancer through early prevention. Smartphone-based Artificial Intelligence (AI) enabled screening of potentially malignant oral lesions offers a scalable solution for this in resource-constrained settings. However, developing accurate and explainable AI segmentation models require high-quality, pixel-level annotated data. This process that is prohibitively expensive, time-consuming, and prone to inter-observer subjectivity among clinical experts. Methods: We designed and empirically validated a Deep Learning-driven Human-in-the-Loop (HITL) framework to pixel-annotate a dataset of 3026 clinical oral images. Using an iterative pseudo-labeling pipeline, we evaluated the model's learning dynamics and performance evolution across five training cycles. We conducted controlled experiments to quantify the networks tolerance to intermediate level of label noise (unreviewed pseudo-labels) to resolve clinical subjectivity using pixel-wise Cohen's Kappa and the STAPLE consensus algorithm. Results: Iterative self-training produced sustained improvements in lesion detection and spatial localization. However, controlled experiments revealed that including even a modest fraction ({approx}10%) of unreviewed pseudo-labels led to a three-to-four-fold increase in training convergence instability and induced a conservative prediction bias that negatively impacted model recall. When measured against multi-expert ground truth, the model's performance converged with the inter-rater reliability ceiling ({kappa} {approx} 0.65), indicating that its predictions fell within the envelope of human agreement. Conclusions: Our findings emphasize that a final expert-driven quality assurance step remains absolutely essential to mitigate training instability, confirmation bias, and clinically unacceptable drops in recall caused by label noise. Overall, this work provides a scalable, empirically validated blueprint for building domain-specific medical imaging datasets in low-resource global health settings, where the dual challenges of annotation cost and inter-observer variability are most acute.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.