Multi-ancestry modeling improves fine-mapping resolution, protein prediction, and discovery for proteome-wide association studies
Krueger, C. J.; Fischer, M.; Rizwan, T.; Kumar, M. M.; Bhargava, S.; Gerszten, R. E.; Taylor, K. D.; Cho, M. H.; Rotter, J. I.; NHLBI TOPMed Consortium, ; Perera, M.; Hu, X.; Manichaikul, A. W.; Im, H. K.; Wheeler, H. E.
Show abstract
Proteomic predictive models are predominantly trained on cis-acting variants in European-ancestry cohorts, limiting power and predictive accuracy in ancestrally diverse populations. We performed cis- and trans-protein quantitative trait locus (pQTL) mapping and developed protein-prediction models using whole-genome sequencing (WGS) and plasma protein levels (Olink) across four ancestry groups from the Trans-omics for Precision Medicine (TOPMed) Multi-Ethnic Study of Atherosclerosis (MESA): European (EUR, n=1270), African (AFR, n=675), Hispanic (HIS, n=642), and Chinese (CHN, n=366), and a combined population (ALL, n=2953). African-ancestry samples demonstrated improved fine-mapping resolution relative to cohort size, yielding significantly smaller cis-credible sets than European-ancestry samples, consistent with shorter linkage disequilibrium (LD) blocks and greater allele frequency diversity in African-ancestry populations. For the first time, we benchmarked fine-mapping models SuSiE, SuShiE, MultiSuSiE, and SuSiEx with multi-ancestral cohorts, revealing a precision-recall tradeoff driven by model assumptions. Comparing protein-prediction models, multivariate adaptive shrinkage (MASHR) and ultimate deconvolution in R (UDR) outperformed elastic net (EN) regression, with trans-pQTL inclusion and fine-mapping improving prediction performance and proteome-wide association study (PWAS) discovery. Applying our models in PWAS of 10 phenotypes, we discovered 68 protein-phenotype associations in All of Us (AoU) that also replicated in Pan-UK Biobank. MASHR and UDR models identified 60% more protein-phenotype associations than EN. Notably, 32 of these associations were not previously reported in the GWAS Catalog. Overall, our study demonstrates the importance of including multiple ancestries in genomic studies to capture the full spectrum of regulatory variation and improve cross-ancestry generalizability.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Enrichment analyses identify shared associations for 25 quantitative traits in over 600,000 individuals from seven diverse ancestries 97%
- Characterizing substructure via mixture modeling in large-scale genetic summary statistics 97%
- Shared components of heritability across genetically correlated traits 96%
Similar papers in this journal
- Proteome-wide Mendelian randomization in global biobank meta-analysis reveals multi-ancestry drug targets for common diseases 97%
- Polymorphic short tandem repeats make widespread contributions to blood and serum traits 97%
- Meta-analysis fine-mapping is often miscalibrated at single-variant resolution 97%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.