IHGAMP: Pan-cancer HRD prediction from routine H&E whole-slide images using foundation models
Ahmad Zafar, S.; Qin, W.; Chengliang, L.; Khan, A. A.; Nazir, A.; Batool, H.; Khalid, F.; Faisal, M. S.
Show abstract
Homologous recombination deficiency (HRD) confers sensitivity to poly (ADP-ribose) polymerase (PARP) inhibitors and platinum-based chemotherapy, representing a critical biomarker for precision oncology across multiple malignancies. Current HRD assessment relies on next-generation sequencing of genomic scar signatures, but specialized infrastructure requirements, high costs, and prolonged turnaround times limit widespread adoption. These barriers restrict access to HRD testing, particularly in resource-constrained settings where the majority of cancer patients receive care. Pan-cancer HRD prediction has been shown, but robustness across histologies and institutions, leak-safe evaluation, and backbone-dependent generalization remain incompletely characterized. Here we show that IHGAMP (Integrative Histopathology-Genomic Analysis for Molecular Phenotyping), a computational framework using vision transformer foundation models, predicts HRD status from H&E images with an AUROC of 0.766 (95% CI 0.727-0.803) on the TCGA held-out test set using OpenCLIP embeddings, and improves to 0.827 with histopathology-pretrained OpenSlideFM embeddings under the same leak-safe protocol. External evaluation on 927 patients (2,718 whole slide images) from seven independent cohorts demonstrated generalization in adenocarcinoma/serous settings (e.g., CPTAC-LUAD AUROC 0.723) and enabled platinum resistance prediction in PTRC-HGSOC (AUROC 0.673), with attenuation in squamous histologies. Systematic comparison of foundation-model embeddings showed that OpenSlideFM outperformed OpenCLIP internally on TCGA (0.827 vs 0.766 AUROC) and improved external generalization in select cohorts (e.g., CPTAC-LUAD), while performance remained attenuated in squamous histologies; TSS-level embedding norm stability across 710 tissue source sites suggested limited site-driven magnitude shifts. Our findings establish that routine histopathology contains morphology associated with HRD that enables moderate, histology-dependent prediction, supporting a potential screening/triage role to prioritize confirmatory molecular testing where appropriate.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Image-Based Consensus Molecular Subtyping in Rectal Cancer Biopsies and Response to Neoadjuvant Chemoradiotherapy 96%
- Predicting the Tumor Microenvironment Composition and Immunotherapy Response in Non-Small Cell Lung Cancer from Digital Histopathology Images 95%
- Generalizing AI-driven Assessment of Immunohistochemistry across Immunostains and Cancer Types: A Universal Immunohistochemistry Analyzer 94%
Similar papers in this journal
- Multiplexed RNA-FISH-guided Laser Capture Microdissection RNA Sequencing Improves Breast Cancer Molecular Subtyping, Prognostic Classification, and Predicts Response to Antibody Drug Conjugates 96%
- Machine learning-based tissue of origin classification for cancer of unknown primary diagnostics using genome-wide mutation features 96%
- Integrative ensemble modelling of cetuximab sensitivity in colorectal cancer PDXs 96%
Similar papers in this journal
- AI-Driven Predictive Biomarker Discovery with Contrastive Learning to Improve Clinical Trial Outcomes 95%
- Evolutionary states and trajectories characterized by distinct pathways stratify ovarian high-grade serous carcinoma patients 94%
- Targeting the mSWI/SNF Complex in POU2F-POU2AF Transcription Factor-Driven Malignancies 92%
Similar papers in this journal
- Non-invasive multi-cancer detection using DNA hypomethylation of LINE-1 retrotransposons 94%
- Tamoxifen Response at Single Cell Resolution in Estrogen Receptor-Positive Primary Human Breast Tumors 93%
- Circulating tumor DNA analysis in advanced urothelial carcinoma: insights from biological analysis and extended clinical follow-up 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.