Back

Moving Beyond Binary Biomarkers: Machine Learning Model Resolves Concurrent and Molecularly Heterogeneous Mismatch Repair and Homologous Recombination Deficiencies in Prostate Cancer

Sharma, K.; Wilson, D. A.; Wang, Y.; Haider, S.; Bhatlapenumarthi, V.; Boiarsky, D.; Cipriaso, J. M.; Yermakov, L.; Dutta, R.; Coleman, I. M.; Bankhead, A.; Jadhav, S. K.; Neupane, B.; Taylor, B. W.; Morrissey, C.; Schweitzer, M. T.; Montgomery, R. B.; Rao, S.; Nevalainen, M. T.; Nelson, A. A.; Antonarakis, E. S.; Pilea, P. G.; Berchuk, J. E.; Zarrabi, K. K.; Pritchard, C.; Kothari, A.; Chen, H.-Z.; George, B.; Kurzrock, R.; Ha, G.; Nelson, P. S.; Banerjee, A.; Auer, P. L.; Kilari, D.; De Sarkar, N.

2025-10-14 genomics
10.1101/2025.10.08.680958 bioRxiv
Show abstract

Current DNA damage repair (DDR) biomarkers employ binary classifications that fail to capture the molecular complexity of tumors with concurrent repair deficiencies. We used genomics analysis to stratify 672 metastatic prostate cancer patients into 11 DDR subgroups, identifying 51 molecular signatures with weighted roles in class identity. We identified a tumor-mutational-burden very-high subset, characterized by 19 mutations/Mb or more, as a molecularly distinct group characterized by preserved genomic integrity and enhanced immunogenicity. Critically, 2.3 percent of tumors exhibited concurrent TMB-High and HRR mutant phenotypes, while 1.5 percent harbored MMR bi-allelic loss without MMRd (mismatch-repair-deficiency) signatures. Clinical validation in 130 patients demonstrated superior immunotherapy responses in tumors with very high TMB levels. We developed CHIMERA DDR, a probabilistic machine learning tool that integrates these 51 genomic features using a nested Random Forest architecture to infer seven clinically relevant DDR subgroups. After negating model overfit concerns, CHIMERA-DDR showed exceptional classification performance (AUCs 0.919-0.999) to accurately detect MMRd and HRR mutant molecular subtypes with or without concurrent DDR deficiencies, resolving admixed phenotypes to enable precision therapeutic stratification beyond binary methods.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.