Back

Homologous recombination deficiency prediction from whole slide images using label refinement and foundation-model benchmarking in ovarian cancer

Shah, N. A.; Sarwar, M.; Ullah, E.

2026-06-30 pathology
10.64898/2026.06.25.734452 bioRxiv
Show abstract

Background: Homologous recombination deficiency (HRD) is clinically imperative in high-grade serous ovarian carcinoma (HGSOC), particularly because of its association with platinum sensitivity and benefit from poly(ADP-ribose) polymerase inhibitor (PARPi) therapy. However, public datasets rarely contain a complete combination of diagnostic haematoxylin and eosin (H&E) whole-slide images (WSIs), validated clinical HRD assay results, genomic scar scores, BRCA1 promoter methylation data, and treatment-response outcomes. This creates a major barrier for computational pathology studies seeking to develop clinically interpretable models of HRD or PARPi response from routine histology. Objective: We performed an exploratory, leakage-controlled computational pathology benchmarking study to evaluate whether H&E WSIs from TCGA-OV contain a measurable morphology-linked signal associated with research-grade molecular HRD labels, and whether label refinement and pathology foundation-model embeddings alter predictive performance. Methods: We assembled a frozen-primary TCGA-OV WSI cohort comprising 717 tissue-section/biospecimen slides from 316 patients. Diagnostic FFPE DX slides were excluded from model selection because of complete patient overlap with the frozen-primary cohort. Two HRD labels were evaluated: an initial mutation-only molecular label based on BRCA/HR-gene mutation evidence, and a refined methylation-enhanced molecular label that additionally incorporated BRCA1 promoter methylation. Feature extraction was performed using ResNet50, UNI, CONCH, Virchow2, Phikon-v2, and UNI2-h encoders. Patient-level attention-based multiple instance learning (ABMIL) was used with patient-as-bag modelling. Evaluation used patient-level grouped 5-fold x 5-repeat stratified cross-validation, with 25 folds total, bootstrap confidence intervals, and patient-level leakage control. Results: The initial mutation-only label classified 78 patients as positive and 238 as negative. The refined methylation-enhanced label recovered 33 additional positives, resulting in 111 positive and 205 negative patients. Patient-level ABMIL using UNI2-h features achieved the strongest performance for the refined label, with AUROC 0.634 (95% CI 0.571-0.698), AUPRC 0.468 (95% CI 0.390-0.562), balanced accuracy 0.597, sensitivity 0.532, specificity 0.663, F1 score 0.494, and Brier score 0.233. The calibrated threshold was 0.512, yielding TN=136, FP=69, FN=52, and TP=59. Comparative models showed lower discrimination, including UNI2-h with the initial label (AUROC 0.628), Phikon-v2 refined (0.582), Virchow2 refined (0.582), CONCH initial (0.587), ResNet50 refined (0.570), and clinical baselines (AUROC 0.54-0.57). Conclusions: TCGA-OV H&E WSIs contain a modest but reproducible morphology-linked signal associated with research-grade molecular HRD status. However, the AUROC around 0.63, absence of clinical HRD assay labels, lack of genomic scar endpoints in the implemented workflow, and absence of PARPi/platinum response targets prevent clinical interpretation. This study should be interpreted as a proof-of-concept benchmarking framework and methodological foundation for future H&E-based predictive modelling in clinically curated PARPi response cohorts.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Modern Pathology
22 papers in training set
Top 0.1%
18.7%
2
npj Digital Medicine
118 papers in training set
Top 0.6%
9.8%
3
Communications Medicine
113 papers in training set
Top 0.2%
8.0%
4
npj Precision Oncology
53 papers in training set
Top 0.1%
7.3%
5
eBioMedicine
183 papers in training set
Top 0.2%
6.3%
50% of probability mass above
6
Nature Communications
5641 papers in training set
Top 33%
3.6%
7
Clinical Cancer Research
64 papers in training set
Top 0.5%
3.6%
8
The Lancet Digital Health
25 papers in training set
Top 0.1%
3.1%
9
npj Breast Cancer
23 papers in training set
Top 0.2%
2.7%
10
PLOS ONE
5266 papers in training set
Top 42%
2.5%
11
Cell Reports Medicine
153 papers in training set
Top 2%
2.1%
12
The Journal of Molecular Diagnostics
39 papers in training set
Top 0.3%
1.9%
13
Genome Medicine
183 papers in training set
Top 3%
1.7%
14
Cancer Research
130 papers in training set
Top 2%
1.5%
15
Scientific Reports
3612 papers in training set
Top 61%
1.3%
16
Journal of Pathology Informatics
15 papers in training set
Top 0.2%
1.1%
17
BMC Cancer
67 papers in training set
Top 2%
1.1%
18
Journal of Clinical Pathology
15 papers in training set
Top 0.3%
1.1%
19
Breast Cancer Research
36 papers in training set
Top 0.5%
1.0%
20
Cancers
213 papers in training set
Top 4%
1.0%
21
Molecular Oncology
55 papers in training set
Top 1%
0.9%
22
Cancer Research Communications
51 papers in training set
Top 2%
0.9%
23
International Journal of Molecular Sciences
494 papers in training set
Top 15%
0.9%
24
JCO Clinical Cancer Informatics
22 papers in training set
Top 0.7%
0.9%
25
iScience
1154 papers in training set
Top 34%
0.9%
26
JAMA Network Open
130 papers in training set
Top 4%
0.9%
27
JCO Precision Oncology
14 papers in training set
Top 0.4%
0.6%
28
Computational and Structural Biotechnology Journal
242 papers in training set
Top 8%
0.6%
29
Frontiers in Bioinformatics
49 papers in training set
Top 2%
0.6%
30
Journal of Experimental & Clinical Cancer Research
25 papers in training set
Top 0.8%
0.6%