Back

Large-scale meta-analysis of over one million individuals reveals the genetic architecture of 127 complex traits in East Asian populations

Jo, J.; Khor, S.-S.; Chu, S.-K.; Ji, Y.; Ueno, K.; Ono, A.; Chen, C.-W.; Do, A.; Han, H.; Kawai, Y.; Kim, N.-E.; Chen, C.-h.; Tokunaga, K.; Won, S.; Yang, H.-C.

2026-06-23 genetics
10.64898/2026.06.18.730290 bioRxiv
Show abstract

Genome-wide association studies (GWASs) have disproportionately focused on European (EUR) populations, limiting the characterization of genetic architecture in other ancestries. To address this imbalance, we integrated large-scale biobanks from Japan, Korea, Taiwan, and China to perform the largest phenome-wide meta-analysis to date in East Asian (EAS) populations, encompassing over one million individuals across 127 complex traits. We identified 8,010 previously unreported associations and observed substantial genetic sharing across EAS subpopulations, while also detecting cohort-specific heterogeneity within the broader EAS context. Transethnic analyses revealed moderate genetic correlations between EAS and EUR populations, indicating both shared and ancestry-specific components of disease risk. Pleiotropy analyses highlighted prominent signals within the HLA region, supported by protein-protein interaction connectivity and immune-related pathway enrichment. Decomposition of genome-wide association matrices further uncovered structured cross-trait architectures, revealing a predominantly shared polygenic backbone driven by metabolic, biochemical, and anthropometric traits, together with two discrete latent components enriched for immune-related processes. Together, our findings refine the genetic architecture of complex traits in East Asian populations at unprecedented scale and clarify the balance between shared and population-specific determinants of human diseases.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.