Back

Cross-ancestry proteome-wide Mendelian randomization prioritizes 12 plasma protein candidates for breast cancer risk

Wu, X.; Godbole, D.; Williams, J.; Sharma, J.; Choi, J.; Liu, Z.; Kraft, P.; Zhang, H.

2026-05-04 epidemiology
10.64898/2026.05.04.26352055 medRxiv
Show abstract

The plasma proteome provides a molecular bridge between genetic variation and disease risk, yet its contribution to breast cancer susceptibility across ancestries remains unclear. We conducted a proteome-wide Mendelian randomization (MR) study of 2,923 plasma proteins using cis-protein quantitative trait loci from 34,557 European participants in the UK Biobank Pharma Proteomics Project, integrated with genome-wide association studies of 156,901 breast cancer cases and 204,634 controls of European, East Asian, and African ancestries. Cross-ancestry meta-analysis identified 12 candidate proteins associated with breast cancer risk (P < 2.5x10-5), including six previously reported and six newly implicated in MR studies. DNPH1 showed cross-ancestry heterogeneity, with a risk-increasing association in European populations and a nominally inverse association in East Asian populations. CASP8, RALB, and USP28 displayed subtype-differentiated associations. Orthogonal validation provided variable support: six demonstrated strong evidence of statistical colocalization; four replicated in an independent European proteomic dataset (deCODE, n = 35,559); two replicated in an independent East Asian proteomic dataset (JCTF, n = 1,384); and four were supported by polygenic-score analyses in the ancestrally diverse All of Us cohort (9,250 cases, 214,857 controls). These findings prioritize a high-confidence subset of plasma proteins, including LRRC25, PARK7, and LRRC37A2, for future mechanistic and translational investigation.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.