Population specific reference panels are crucial for the genetic analyses of Native Hawai’ians: an example of the CREBRF locus
Lin, M.; Caberto, C.; Wan, P.; Li, Y.; Lum-Jones, A.; Tiirikainen, M.; Pooler, L.; Nakamura, B.; Sheng, X.; Porcel, J.; Lim, U.; Setiawa, V. W.; Le Marchand, L.; Wilkens, L. R.; Haiman, C. A.; Cheng, I.; Chiang, C. W. K.
Show abstract
Statistical imputation applied to genome-wide array data is the most cost-effective approach to complete the catalog of genetic variation in a study population. However, imputed genotypes in underrepresented populations incur greater inaccuracies due to ascertainment bias and a lack of representation among reference individuals,, further contributing to the obstacles to study these populations. Here we examined the consequences due to the lack of representation by genotyping a functionally important, Polynesian-specific variant, rs373863828, in the CREBRF gene, in a large number of self-reported Native Hawaiians (N=3,693) from the Multiethnic Cohort. We found the derived allele of rs373863828 was significantly associated with several adiposity traits with large effects (e.g. 0.214 s.d., or approximately 1.28 kg/m2, per allele, in BMI as the most significant; P = 7.5x10-5). Due to the current absence of Polynesian representation in publicly accessible reference sequences, rs373863828 or any of its proxies could not be tested through imputation using these existing resources. Moreover, the association signals at this Polynesian-specific variant could not be captured by alternative approaches, such as admixture mapping. In contrast, highly accurate imputation can be achieved even if a small number (<200) of Polynesian reference individuals were available. By constructing an internal set of Polynesian reference individuals, we were able to increase sample size for analysis up to 3,936 individuals, and improved the statistical evidence of association (e.g. p = 1.5x10-7, 3x10-6, and 1.4x10-4 for BMI, hip circumference, and T2D, respectively). Taken together, our results suggest the alarming possibility that lack of representation in reference panels would inhibit discovery of functionally important, population-specific loci such as CREBRF. Yet, they could be easily detected and prioritized with improved representation of diverse populations in sequencing studies.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Leveraging phenotypic variability to identify genetic interactions in human phenotypes 97%
- Characterization of exome variants and their metabolic impact in 6,716 American Indians from Southwest US 97%
- Enrichment analyses identify shared associations for 25 quantitative traits in over 600,000 individuals from seven diverse ancestries 96%
Similar papers in this journal
- Whole Genome Sequencing Analysis Of Body Mass Index Identifies Novel African Ancestry-Specific Risk Allele 97%
- Comprehensive genomic analysis of dietary habits in UK Biobank identifies hundreds of genetic loci and establishes causal relationships between educational attainment and healthy eating 97%
- Quantifying the contribution of Neanderthal introgression to the heritability of complex traits 96%
Similar papers in this journal
- Ancestral diversity improves discovery and fine-mapping of genetic loci for anthropometric traits - the Hispanic/Latino Anthropometry Consortium 97%
- Multivariate adaptive shrinkage improves cross-population transcriptome prediction for transcriptome-wide association studies in underrepresented populations 96%
- Polygenic risk score prediction accuracy convergence 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.