Back

PRECISE: Benchmarking digital pathology with expert-annotated contiguous IHC-H&E serial prostate sections

Calapaqui Teran, A. K.; Gonzalez Bernad, A. A.; Cobo Cano, M.; Sanchez Magdaleno, L.; Marcos Gonzalez, S.; Delgado Bolton, R. C.; Moustafa Calvo, J.; Gomez Roman, J. J.; Lara, L.

2026-07-22 pathology
10.64898/2026.07.21.26358559 medRxiv
Show abstract

We present PRECISE (PRostate Expert-annotated Contiguous IHC-H\&E Serial sEctions), a hybrid histopathology dataset of paired hematoxylin and eosin (H\&E) and immunohistochemistry (IHC) whole-slide images (WSIs), comprising 37 prostate core needle biopsies from 25 patients, each with matched H\&E and CKAPM+racemase staining. To the best of our knowledge, this is the first publicly available dataset offering spatially harmonized, pixel-level expert annotations across both staining modalities in prostate biopsy WSIs - directly mirroring the two-stage (H\&E-then-IHC) clinical diagnostic workflow used to resolve morphological uncertainty, restricted to cases in which that workflow reached diagnostic consensus. The dataset contains 24,387 annotations spanning seven diagnostically critical classes: malignant glands, benign glands, stromal tissue, intraductal carcinoma (IDC-P), high-grade prostatic intraepithelial neoplasia (HGPIN), atypical intraductal proliferation (AIP), and tissue artifacts. Unlike existing resources, which focus on binary tumor classification or lack IHC pairing, this dataset captures the full morphological spectrum encountered in routine prostate pathology, including rare precursor lesions and confounding entities underrepresented in current benchmarks. Annotations were validated through a structured three-stage consensus by two expert uropathologists, with IHC serving as biological ground truth for boundary definition. PRECISE is designed as a robust benchmark for multimodal semantic segmentation and self-supervised learning, and is openly released to promote reproducible research and accelerate AI-assisted diagnosis in prostate cancer.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.