Back

Diverse Genomes, Shared Health: Insights from a Health System Biobank

Haas, R.; Margolis, M. P.; Wei, A.; Yamaguchi, T. N.; Feng, J.; Tran, T.; Tozzo, V.; Queen, K. J.; Mootor, M. F. E.; Patil, V.; Broudy, M. E.; Tung, P.; Alam, S.; Martinez, D. B.; Patel, Y.; Zeltser, N.; Hugh-White, R.; Arbet, J.; Caggiano, C.; Shemirani, R.; Tian, M.; Thapaliya, P.; Eloyan, L.; Chen, L. O.; Ariannejad, M.; Lajonchere, C.; UCLA Precision Health Data Discovery Repository Working Group, ; UCLA Precision Health ATLAS Working Group, ; UCLA Health IT HPC Team, ; Regeneron Genetics Center, ; Pasaniuc, B.; Bui, A.; Arboleda, V. A.; Chang, T. S.; Zaitlen, N.; Spellman, P. T.; Bout

2025-06-12 genetic and genomic medicine
10.1101/2025.06.11.25329386 medRxiv
Show abstract

Linking genetic data with electronic health records in hospital biobanks promises to advance precision medicine, but limited ancestral diversity constrains discovery and generalizability. We analyzed 93,936 participants from the UCLA ATLAS Community Health Initiative to inform disease prevalence and genetic risk across five continental and 36 fine-scale ancestry groups. We discovered numerous unreported gene-phenotype associations, including FN3K with intestinal disaccharidase deficiency in Europeans and admixed Americans. Polygenic scores (PGS) robustly predicted common diseases, with effects markedly diminished in non-Europeans. Furthermore, we reduced the pronounced European bias in curated clinical variants using computational predictors, uncovering unreported disease-gene associations, including ANKZF1 and peripheral vascular disease in AFR. Longitudinal data revealed that semaglutide efficacy varies across ancestries, is associated with PGS for type 2 diabetes, and is modulated by genetic variation in PTPRU. These findings illustrate how ancestrally diverse biobanks from a single health system yield robust disease associations and pharmacogenomic insights.

Published in Cell (predicted rank #6) · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.