The Biobank Rare Variant consortium powers the discovery of rare genetic associations through global collaboration
Palmer, D. S.; Hill, B.; Hodgson, S.; Joeloo, M.; Kalantzis, G.; Kousathanas, A.; Koyama, S.; Lu, W.; Namba, S.; Rodriguez, Z. B.; Shortt, J. A.; Sonehara, K.; Vartanian, N.; Vy, H. M. T.; Wade, I. A.; White, S. L.; Baya, N. A.; Chami, N.; Do, R.; Estrada, K.; Finer, S.; Genovese, G.; Guez, J.; Itan, Y.; Kanai, M.; Lassen, F. H.; Matsuda, K.; Moutsianas, L.; Peloso, G. M.; Priit, P.; Rader, D. J.; Rendon, A.; Rocheleau, G.; Sadeghi-Alavijeh, O.; Selvaraj, M. S.; Smit, R. A.; Wang, D.; Wigdor, E. M.; Yu, Z.; Colorado Center for Personalized Medicine, ; Estonian Biobank Research Team, ; Genes
Show abstract
Rare coding variants can have large effects on disease risk and provide direct routes from human genetics to disease mechanisms and therapeutic targets, but their discovery is constrained by sample size, particularly for low-prevalence diseases. Here we establish the Biobank Rare Variant Analysis (BRaVa) consortium, a global rare variant association resource that integrates sequencing and linked health-record data from ten biobanks and cohorts comprising over 1.2 million individuals across diverse ancestries. We performed gene-based meta-analyses of rare coding variation across 33 clinical endpoints and 11 quantitative traits. Aggregating evidence across biobanks and ancestries identified 514 gene-trait associations, including 31 not previously reported in prior studies or curated association resources following systematic literature review. Notably, 36.1% of gene-level associations were undetectable in any individual biobank, and 91 emerged only through cross-ancestry meta-analysis, demonstrating that federated integration enables discovery beyond the reach of single cohorts. Similar gains were observed at the variant level, where 25.0% of phenotype-locus associations were detectable only through meta-analysis. Effect size estimates were correlated across ancestries with concordant directions of effect, supporting the generalizability of rare variant associations. The identified signals implicate pathways involved in transcriptional and epigenetic regulation, metabolism, vascular and epithelial biology, and immune function, highlighting rare coding variation as an engine for biological discovery across medical record phenotypes. For example, damaging variation in ANKRD12 implicates inflammatory transcriptional dysregulation in asthma and chronic obstructive pulmonary disease, and ultra-rare predicted loss-of-function variants in NAA15 link protein acetylation processes to type 2 diabetes risk. BRaVa establishes a scalable framework and freely available community resource for rare variant meta-analysis across global biobanks. Public release of gene- and variant-level association summary statistics provides a reference map of rare coding variant associations to support disease gene discovery, biological interpretation, and therapeutic target prioritization as sequencing-linked health-record resources continue to expand.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Meta-analysis fine-mapping is often miscalibrated at single-variant resolution 98%
- Proteome-wide Mendelian randomization in global biobank meta-analysis reveals multi-ancestry drug targets for common diseases 97%
- Cystatin C is glucocorticoid-responsive, directs recruitment of Trem2+ macrophages and predicts failure of cancer immunotherapy 97%
Similar papers in this journal
Similar papers in this journal
- Enrichment analyses identify shared associations for 25 quantitative traits in over 600,000 individuals from seven diverse ancestries 97%
- Interaction molecular QTL mapping discovers cellular and environmental modifiers of genetic regulatory effects 97%
- Characterizing substructure via mixture modeling in large-scale genetic summary statistics 97%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.