IdentifiHR: predicting homologous recombination deficiency in high-grade serous ovarian carcinoma using gene expression
Weir, A. L.; Lee, S. C.; Li, M.; Tan, C. W.; Ramus, S. J.; Davidson, N. M.
Show abstract
BackgroundApproximately half of all high-grade serous ovarian carcinomas (HGSCs) have a therapeutically targetable defect in the homologous recombination (HR) DNA repair mechanism. While there are genomic and transcriptomic methods, developed for other cancer types, to identify HR deficient (HRD) samples, there are no gene expression-based tools to predict HR repair status in HGSC specifically. We have built the first HGSC-specific model to predict HR repair status using gene expression. MethodsWe separated The Cancer Genome Atlas (TCGA) cohort of HGSCs (n = 361) into training (n = 288) and testing (n = 73) sets and labelled each case as HRD or HR proficient (HRP) based on the clinical standard for classification, being a score of HRD genomic damage. Using the training set, we performed differential gene expression analysis between HRD and HRP cases. The 2604 significantly differentially expressed genes were then used to tune and train a penalised logistic regression model. ResultsIdentifiHR is an elastic net penalised logistic regression model that uses the expression of 209 genes to predict HR status in HGSC. These genes capture known regions of HR-specific copy number alteration, which impact gene expression levels, and preserve the genomic damage signal. IdentifiHR has an accuracy of 85% in the TCGA test set and of 91% in an independent cohort of 99 samples, collected from primary tumours before (n = 74/99) and after autopsy (n = 6/99), in addition to ascites (n = 12/99) and normal fallopian tube samples (n = 7/99). Further, IdentifiHR is 84% accurate in pseudobulked single-cell HGSC sequencing from 37 patients and outperforms existing gene expression-based methods to predict HR status, being BRCAness, MutliscaleHRD and expHRD. ConclusionsIdentifiHR is an accurate model to predict HR status in HGSC using gene expression alone, that is available as an R package from https://github.com/DavidsonGroup/IdentifiHR.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Multi-scale characterisation of homologous recombination deficiency in breast cancer 97%
- Genomic analysis of patient-derived xenograft models reveals intra-tumor heterogeneity in endometrial cancer and can predict tumor growth inhibition with talazoparib 96%
- Evolving neoantigen profiles in colorectal cancers with DNA repair defects 95%
Similar papers in this journal
- Pan-cancer Analysis of Homologous Recombination Deficiency in Cell Lines 95%
- Drug-gene interaction screens coupled to tumour data analyses identify the most clinically-relevant cancer vulnerabilities driving sensitivity to PARP inhibition 95%
- Predicting tumor immune microenvironment and checkpoint therapy response of head & neck cancer patients from blood immune single-cell transcriptomics 93%
Similar papers in this journal
Similar papers in this journal
- Acquired RAD51C promoter methylation loss causes PARP inhibitor resistance in high grade serous ovarian carcinoma 95%
- Aberrant transcript usage induces homologous recombination deficiency and predicts therapeutic responses 94%
- Germline and somatic genetic variants in the p53 pathway interact to affect cancer risk, progression and drug response 94%
Similar papers in this journal
- Patient-derived organoids identify tailored therapeutic options and determinants of plasticity in sarcomatoid urothelial bladder cancer 95%
- Tumor break load quantitates structural variant-associated genomic instability with biological and clinical relevance across cancers 94%
- Single Cell RNA-sequencing of BCG naive and recurrent non-muscle invasive bladder cancer reveals a CD6/ALCAM mediated immune-suppressive pathway 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.