Back

Integrated Clinicogenomic Risk Modeling for Metachronous Second Primary Cancers

Amsalem, J.; Ostrovnaya, I.; Marderstein, A. R.; Liu, Y. L.; Perea-Chamblee, T.; Ravichandran, V.; Jee, J.; Conry, M.; Khurram, A.; Kemel, Y.; Kim, E.; Mukherjee, S.; Latham, A.; Banaszak, L.; Kundra, R.; Magunta, S.; Fong, C.; Buas, M. F.; Bandlamudi, C.; Bernstein, J.; Seshan, V.; Salles, G.; Mandelker, D.; Berger, M. F.; Solit, D. B.; Stadler, Z. K.; Carrot-Zhang, J.; Schultz, N.; Offit, K.; Joseph, V.

2026-06-22 genetic and genomic medicine
10.64898/2026.06.12.26355388 medRxiv
Show abstract

Improvements in cancer survival have increased the burden of subsequent primary malignancies. We developed and validated a programmatic classifier of multiple primary cancers (MPC) to derive second cancer phenotypes at scale. Among 81,175 cancer patients, we identified 56 first-second cancer pairs, 22 of which exceeded SEER primary cancer incidence rates. Even after accounting for various known risk factors, substantial elevated risk persisted, even in established hereditary cancer pairs (breast-ovary, breast-pancreas, prostate-pancreas), suggesting that current screening protocols do not adequately account for MPC susceptibility. To address this limitation, we built machine-learning models integrating rare germline variants, polygenic risk scores, treatment exposures, and demographic features to predict site-specific second primaries in breast and prostate cancer survivors. These models accurately predicted second ovarian and pancreatic cancers across a long follow-up period (15-year time-dependent AUC 0.70). This is the first systematic, pan-cancer integration of clinicogenomic factors for early prediction of second-primary malignancies. Our framework enables individualized risk estimation, enhanced targeted surveillance, and cancer prevention amongst a growing population of cancer survivors.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.