Back

Assessing the risk stratification of breast cancer polygenic risk scores in two Brazilian samples

Barreiro, R. A.; Almeida, T. F.; Gomes, C. S.; Monfardini, F.; Farias, A. A.; Tunes, G. C.; Souza, G. M.; Duim, E.; Correia, J. S.; Coelho, A. V.; Caraciolo, M. P.; Duarte, Y. A.; Zatz, M.; Amaro, E.; Oliveira, J. B.; Bitarello, B. D.; Brentani, H.; Naslavsky, M. S.

2022-09-10 genetic and genomic medicine
10.1101/2022.09.09.22279721 medRxiv
Show abstract

Polygenic risk scores (PRS) for breast cancer (BC) have a clear clinical utility in risk prediction. PRS transferability across populations and ancestry groups is hampered by population-specific factors, ultimately leading to differences in variant effects, such as linkage disequilibrium (LD) and differences in variant frequency (AF-diff). Thus, locally-sourced population-based phenotypic and genomic datasets are essential to assess the validity of PRS derived from signals detected across populations. Here, assess the transferability of a BC PRS composed of 313 risk variants (313-PRS) in two Brazilian tri-hybrid admixed ancestries (European, African and Native American) whole-genome sequenced cohorts. We computed 313-PRS in both cohorts (n=753 and n=853) versus the UK Biobank (UKBB, n=264,307) as reference. We show that although the Brazilian cohorts have a high European (EA) component, with AF-diff and to a lesser extent LD patterns like those found in EA populations, the 313-PRS distribution is inflated when compared to that of the UKBB, leading to potential overestimation of PRS-based risk if EA is taken as a standard. Interestingly, we find that case-controls lead to equivalent predictive power when compared to UKBB-EA samples with AUROC values of 0.66-0.62 compared to 0.63 for UKBB.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.