Human Genetics and Genomics Advances
○ Elsevier BV
All preprints, ranked by how well they match Human Genetics and Genomics Advances's content profile, based on 84 papers previously published here. The average preprint has a 0.08% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Carlson, J. C.; Krishnan, M.; Liu, S.; Anderson, K. J.; Zhang, J. Z.; Yapp, T.-A. J.; Chiyka, E. A.; Dikec, D. A.; Cheng, H.; Naseri, T.; Reupena, M. S.; Viali, S.; Deka, R.; Hawley, N. L.; McGarvey, S. T.; Weeks, D. E.; Minster, R. L.
Show abstract
Genotype imputation is fundamental to association studies, and yet even gold standard panels like TOPMed are limited in the populations for which they yield good imputation. Specifically, Pacific Islanders are poorly represented in extant panels. To address this, we constructed an imputation reference panel using 1,285 Samoan individuals with whole-genome sequencing, combined with 1000 Genomes Project (1KGP) individuals, to create a reference panel that better represents Pacific Islander, specifically Samoan, genetic variation. We compared this panel to 1KGP and TOPMed-R3 panels based on imputed variants using genotyping array data for 1,834 Samoan participants who were not part of the panels. The 1KGP + 1285 Samoan panel yielded up to two times more well-imputed (r2 [≥] 0.80) variants than TOPMed-R3 and 1KGP and was enriched for moderate and high impact variants. There was improved imputation accuracy across the minor allele frequency (MAF) spectrum, although it was most pronounced for variants with 0.01 [≤] MAF [≤] 0.05. Imputation accuracy (r2) was greater for population-specific variants (high fixation index, FST) and those from larger haplotypes (high LD score). However, the gain in imputation accuracy over TOPMed-R3 was largest for small haplotypes (low LD score), reflecting the Samoan panels ability to capture population-specific variation not well tagged by other panels. We also augmented the 1KGP reference panel with varying numbers of Samoan participants and found that panels with 24 Samoans yielded similar performance to TOPMed-R3, and panels with 48 or more Samoans included outperformed TOPMed-R3 for all variants with MAF [≥] 0.001. Meta imputation of the TOPMed-R3 and 1285 Samoan panels yielded poorer performance than the Samoan only panel. We also demonstrated that the phasing of the reference panel impacts the imputation of population-specific variants when the reference panel is composed of individuals from an isolated population and not combined with ancestrally diverse haplotypes. This study identifies variants with improved imputation using population-specific reference panels and provides a framework for constructing other population-specific reference panels.
Movassagh, M.; Newbury, L.; Hehnly, C.; Whalen, A.; Peterson, M.; Mondragon Estrada, E.; Ericson, J.; Smith, J.; Sasanami, M.; Natukwatsa, D.; Mugamba, J.; Ssenyonga, P.; Onen, J.; Burgoine, K.; Zhang, L.; Olupot-Olupot, P.; Kumbakumba, E.; Wegoye, E.; Ochora, M.; Mulondo, R.; Mbabazi-Kabachelor, E.; Fronterre, C.; Broach, J.; Paulson, J.; Morton, S.; Schiff, S.
Show abstract
BackgroundNeonatal disorders such as post-infectious hydrocephalus exhibit a higher incidence in Africa, where the intricate relationships between genetic ancestry, environmental exposures, and other risk factors likely contribute to the increased incidence. MethodsTo start to characterize the common genetic architecture of Ugandan infants, we analyzed genome sequencing data from 1,030 Ugandan infants recruited from studies targeting neonatal sepsis and hydrocephalus. We employed genetic admixture analysis and integrated geospatial data to examine the relationships between genetic backgrounds and disease prevalence within this cohort. ResultsOur results identified four distinct genetic admixture groups, each correlating strongly with specific geographic distributions across Uganda. Notably, a predominance of one admixture group, most common in northern Uganda, was overrepresented in the participants with post-infectious hydrocephalus. ConclusionThis study underscores the importance of genetic factors in disease manifestation at the population level, and a role for such precision public health approaches in complex neonatal disorders in African populations.
Yapp, T.-A. J.; Krishnan, M.; Liu, S.; Manna, S. L.; Cheng, H.; Naseri, T.; Reupena, M. S.; Viali, S.; Tuitele, J.; Deka, R.; Hawley, N. L.; McGarvey, S. T.; Weeks, D. E.; Minster, R. L.; Carlson, J. C.
Show abstract
Dyslipidemia is a significant risk factor for cardiovascular disease (CVD), the leading cause of death in Samoa, accounting for 34% of deaths. Polygenic scores (PGS) derived from large scale multi ancestry genome-wide association studies offer potential for improved CVD risk prediction by aggregating genetic effects on lipid traits, yet their performance in Pacific Islander populations remains largely unknown. We evaluated the transferability of multi-ancestry PGS for LDL cholesterol (LDL C), HDL cholesterol (HDL C), triglycerides (TG), and total cholesterol (TC) in 4,342 Samoan adults across five cohorts spanning 1990 to 2010. PGS derived from Graham et al. and Kanoni et al. multi-ancestry meta-analyses were harmonized with genome-wide imputed genotypes using a Samoan-specific reference panel, and performance was assessed using incremental R^2 from linear mixed models with bootstrapped confidence intervals. PGS performance varied across traits and cohorts: HDL C showed the highest performance (incremental R^2 5.0 to15.0%), followed by LDL C (5.7 to 8.6%) and TC (5.0 to10.7%), with TG showing the lowest performance (3.5 to 7.0%). Meaningful LDL C transferability was achieved only when using a genome-wide PRS CS score (99.6 to 99.7% variant matching), whereas a curated pruning-and-thresholding score achieved only ~9% matching and near-zero performance. These findings establish the first systematic benchmarks for lipid PGS performance in Samoans, demonstrate that multi-ancestry scores can achieve meaningful transferability in this underrepresented population when genome-wide variant coverage is ensured, and highlight the importance of rigorous variant harmonization assessment prior to clinical deployment of PGS in diverse populations.
Barnett, E. J.; Hess, J. L.; Hou, J.; Escott-Price, V.; Fennema-Notestine, C.; Kremen, W.; Lin, S.-J.; Zhang, C.; Gaiteri, C.; Elman, J.; Holmans, P.; Faraone, S. V.; Glatt, S. J.
Show abstract
Genetic risk factors for neuropsychiatric disorders are well documented. However, some individuals with high genetic risk remain unaffected, and the mechanisms underlying such resilience remain poorly understood. The presence of protective resilience factors that mitigate risk could help explain the disconnect between predicted risk and reality, particularly when genetic contributions are substantial but incompletely understood. Identifying and studying resilience factors could improve our understanding of pathology, enhance risk prediction, and inform preventive measures or treatment strategies. However, such efforts are complicated by the difficulty of identifying resilience that is separable from low risk. We developed a novel adversarial multi-task neural network model to detect genetic resilience markers. The model learns to separate high-risk unaffected individuals from affected individuals at similar risk while "unlearning" patterns found in low-risk groups using adversarial learning. In simulated and existing Alzheimers disease (AD) datasets, we identified markers of resilience with a feature-importance-based approach that prioritized specificity, generated resilience scores, and analyzed associations with polygenic risk scores (PRS). In simulations, our model had high specificity and sensitivity in identifying resilience markers, significantly outperforming traditional approaches. Applied to AD data, the model generated genetic resilience scores protective against AD and independent of PRS. We identified five resilience-associated SNPs, including known AD-associated variants, underscoring their potential involvement in resilience. Our findings support the utility of resilience scores in modifying risk predictions, particularly for high-risk groups. Expanding this method could aid in understanding resilience mechanisms, potentially improving diagnosis, prevention, and treatment strategies for AD and other brain disorders.
Lo, Y.-C.; Chan, T. F.; Jeon, S.; Maskarinec, G.; Taparra, K.; Nakatsuka, N.; Yu, M.; Chen, C.-Y.; Lin, Y.-F.; Wilkens, L. R.; Le Marchand, L.; Haiman, C. A.; Chiang, C. W. K.
Show abstract
Polygenic scores (PGS) are promising in stratifying individuals based on the genetic susceptibility to complex diseases or traits. However, the accuracy of PGS models, typically trained in European- or East Asian-ancestry populations, tend to perform poorly in other ethnic minority populations, and their accuracies have not been evaluated for Native Hawaiians. Using body mass index, height, and type-2 diabetes as examples of highly polygenic traits, we evaluated the prediction accuracies of PGS models in a large Native Hawaiian sample from the Multiethnic Cohort with up to 5,300 individuals. We evaluated both publicly available PGS models or genome-wide PGS models trained in this study using the largest available GWAS. We found evidence of lowered prediction accuracies for the PGS models in some cases, particularly for height. We also found that using the Native Hawaiian samples as an optimization cohort during training did not consistently improve PGS performance. Moreover, even the best performing PGS models among Native Hawaiians would have lowered prediction accuracy among the subset of individuals most enriched with Polynesian ancestry. Our findings indicate that factors such as admixture histories, sample size and diversity in GWAS can influence PGS performance for complex traits among Native Hawaiian samples. This study provides an initial survey of PGS performance among Native Hawaiians and exposes the current gaps and challenges associated with improving polygenic prediction models for underrepresented minority populations.
Wendt, F. R.; Pathak, G. A.; Vahey, J.; Qin, X.; Koller, D.; Cabrera-Mendoza, B.; Haeny, A.; Harrington, K. M.; Rajeevan, N.; Duong, L. M.; Levey, D. F.; De Angelis, F.; De Lillo, A.; Bigdeli, T. B.; Pyarajan, S.; VA Million Veteran Program, ; Gaziano, J. M.; Gelernter, J.; Aslan, M.; Provenzale, D.; Helmer, D. A.; Hauser, E. R.; Polimanti, R.; Department of Veteran Affairs Cooperative Study Program (#2006),
Show abstract
The Million Veteran Program (MVP) participants represent 100 years of US history, including significant social and demographic change over time. Our study assessed two aspects of the MVP: (i) longitudinal changes in population diversity and (ii) how these changes can be accounted for in genome-wide association studies (GWAS). The MVP was divided into five birth cohorts (N-range=123,888 [born from 1943-1947] to 136,699 [born from 1948-1953]). Groups of participants were defined by (i) HARE (harmonized ancestry and race/ethnicity) and (ii) a random-forest clustering approach using the 1000 Genomes Project and the Human Genome Diversity Project (1kGP+HGDP) reference panels (77 world populations representing six continental groups). In these groups, we performed GWASs of height, a trait potentially affected by population stratification. Birth cohorts demonstrate important trends in ancestry diversity over time. More recent HARE-assigned Europeans, Africans, and Hispanics had lower European ancestry proportions than older birth cohorts (0.010<Cohens d<0.259, p<7.80x10-4). Conversely, HARE-assigned East Asians showed an increase in European ancestry proportion over time. In GWAS of height using HARE assignments, genomic inflation due to population stratification was prevalent across all birth cohorts (linkage disequilibrium score regression intercept=1.08{+/-}0.042). The 1kGP+HGDP-based ancestry assignment significantly reduced the population stratification (mean intercept reduction=0.045{+/-}0.007, p<0.05) confounding in the GWAS statistics. This study provides a comprehensive characterization of ancestry diversity of the MVP cohort over time and highlights that more refined modeling of genetic diversity (e.g., the 1kGP+HGDP-based ancestry assignment) can more accurately capture the polygenic architecture of traits and diseases that could be affected by population stratification.
Spor, L. M.; Liau, E. M.; Sanchis-Juan, A.; Silva, A. N.; Cheng, H.; Naseri, T.; Reupena, M. S.; Viali, S.; Tuitele, J.; Kershaw, E. E.; Deka, R. D.; Hawley, N. L.; McGarvey, S. T.; Weeks, D. E.; Carlson, J. C.; Brand, H.; Minster, R. L.
Show abstract
Structural variants (SVs) are often excluded from genetic research because they are difficult to call, but they can have substantial effects on phenotypic traits. SVs have not previously been characterized in Samoans, an understudied population with a high burden of complex diseases. Using short-read whole genome sequencing data, we called SVs in 1,276 Samoans and created a Samoan-specific imputation panel inclusive of both SVs and single nucleotide variants (SNVs), called the Soifua Manuia-SV panel. Using this panel, we imputed SVs and SNVs in 3,611 Samoans with array data, enabling analysis of SV-phenotype associations in a sample of 4,887 Samoan participants. We evaluated imputation performance in Samoans against two other reference panels: (i) an SNV-only Samoan-specific reference panel, to assess whether SV inclusion impacts SNV imputation, and (ii) an SV and SNV, multi-ancestry reference panel composed of 1000 Genomes participants, which did not include Polynesians, to assess the importance of including the target population in the reference panel. The Soifua Manuia-SV panel substantially outperformed the multi-ancestry SV and SNV panel, yielding 5.5 million more high-quality (r2[≥]0.8) variants, including over 8,000 more high-quality SVs. SNV imputation based on the two Samoan-specific panels performed similarly overall, suggesting that SV inclusion does not strongly impact SNV imputation quality. This work highlights the importance of population representation for accurate imputation.
Lambert, S. A.; Wingfield, B.; Gibson, J. T.; Gil, L.; Ramachandran, S.; Yvon, F.; Saverimuttu, S.; Tinsley, E.; Lewis, E.; Ritchie, S. C.; Wu, J.; Canovas, R.; McMahon, A.; Harris, L. W.; Parkinson, H.; Inouye, M.
Show abstract
Polygenic scores (PGS) have transformed human genetic research and have multiple potential clinical applications, including risk stratification for disease prevention and prediction of treatment response. Here, we present a series of recent enhancements to the PGS Catalog (www.PGSCatalog.org), the largest findable, accessible, interoperable, and reusable (FAIR) repository of PGS. These include expansions in data content and ancestral diversity as well as the addition of new features. We further present the PGS Catalog Calculator (pgsc_calc, https://github.com/PGScatalog/pgsc_calc), an open-source, scalable and portable pipeline to reproducibly calculate PGS that securely democratizes equitable PGS applications by implementing genetic ancestry estimation and score normalization using reference data. With the PGS Catalog & calculator users can now quantify an individuals genetic predisposition for hundreds of common diseases and clinically relevant traits. Taken together, these updates and tools facilitate the next generation of PGS, thus lowering barriers to the clinical studies necessary to identify where PGS may be integrated into clinical practice.
Lancaster, H. S.; Dinu, V.; Li, J.; Gruen, J. R.; GRaD Consortium,
Show abstract
PurposeReading ability is a complex skill utilizing multiple proficiencies and that develops through interactions between genetic and environmental factors. This study presents an alternative analytic pipeline to identify key genetic and demographic contributors to reading ability. MethodsWe analyzed data from the Avon Longitudinal Study of Parents and Children (ALSPAC; N = 3 232) using a multi-step analytical pipeline. To reduce measurement error, we generated a latent reading ability score. We selected single nucleotide polymorphisms (SNPs) based on existing literature and genome-wide association studies (GWAS). We applied elastic net regression to identify informative predictors in two models, a SNP-only model and a SNP-plus demographic, environmental, and behavioral variables model. We compared the SNP-based heritability estimates and R2 from the fitted models. We also performed pathway enrichment analysis on the informative SNPs. ResultsThe traditional GWAS identified one genome-wide significant SNP on chromosome X and produced a moderate heritability estimate of .23 (SE = 0.07). We included 148 SNPs in the elastic net models. The SNP-only model identified 61 informative SNPs (R2 = .12), whereas the SNP-plus model identified 96 informative SNPs (R2 = .32). The SNP-plus model also showed that several behavioral characteristics positively predicted latent reading ability. Enrichment analysis revealed overrepresentation of several biological pathways among the informative SNPs. ConclusionsThis study shows that our analytic pipeline can identify important genetic and demographic predictors of reading ability, providing a powerful alternative to traditional methods and contributing to a deeper understanding of the factors that drive reading development.
Cataldo-Ramirez, C.; Lin, M.; McMahon, A.; Gignoux, C.; Weaver, T. D.; Henn, B. M.
Show abstract
Genome-wide association studies (GWAS) and polygenic score (PGS) development are typically constrained by the data available in biobank repositories in which European cohorts are vastly overrepresented. Here, we increase the utility of non-European participant data within the UK Biobank (UKB) by characterizing the genetic affinities of UKB participants who self-identify as Bangladeshi, Indian, Pakistani, "White and Asian" (WA), and "Any Other Asian" (AOA), towards creating a more robust South Asian sample size for future genetic analyses. We assess the relationships between genetic structure and self-selected ethnic identities and use consistent patterns of clustering in the dataset to train a support vector machine (SVM). The SVM was utilized to reassign n = 1,853 AOA and WA participants at the subcontinental level, and increase the sample size of the UKB South Asian group by 1,381 additional participants. We further leverage these samples to assess GWAS performance and PGS development. We include environmental covariates in the height GWAS by implementing a rigorous covariate selection procedure, and compare the outputs of two GWAS models: GWASnull and GWASenv. We show that PGS performance derived from both GWAS models yield comparable prediction to PGS models developed with an order of magnitude larger training, and environmentally-adjusted PGS models reduce the sex-bias in predictive performance. In summary, we demonstrate how GWAS performance can be improved by leveraging ambiguous ethnicity codes, ancestry matched imputation panels, and including environmental covariates.
Cordogan, S.; Starr, D. B.; Treff, N. R.; Lanchbury, J.; Goldstein, E.; Sadeghi, K.; Burmeister, M.; Dayani, L.; Macdonald, P.; Keen-Kim, J. D.; Fishel, S.; Cervantes, E.; Folkersen, L.
Show abstract
Polygenic risk scores (PRSs) can reduce lifetime disease risk by guiding embryo selection during in vitro fertilization (IVF). We performed genome-wide association meta-analyses totaling [~]1.5 million individuals to construct state-of-the-art PRSs for nine diseases: Alzheimers disease, breast cancer, coronary artery disease, endometriosis, hypertension, prostate cancer, rheumatoid arthritis, type 1 diabetes, and type 2 diabetes. The resulting predictors achieved liability-scale R2 values of up to 22.9% for type 2 diabetes, matching or exceeding previously published benchmarks across all scores. Three PRS - Alzheimers disease, prostate cancer, and type 2 diabetes - explained over 75% of the common-SNP heritability. Within-family validation in 40,872 siblings across 18,840 families showed that, for eight of nine diseases, predictive performance was comparable to population-level results, confirming substantial direct genetic effects. Modeling of embryo selection suggests that couples with five euploid embryos could achieve 27-67% relative risk reduction across diseases. While the limitations of genetic data availability meant that these estimates were performed in European ancestry samples, we performed validation in the multi-ancestry US-based All of Us Biobank, demonstrating significant statistical power across ancestries. These findings support the clinical applicability of PRS-guided embryo selection to reduce the burden of common diseases.
Martinez, K. L.; Klein, A.; Martin, J. R.; Sampson, C. U.; Giles, J. B.; Beck, M. L.; Bhakta, K.; Quatraro, G.; Farol, J.; Karnes, J. H.
Show abstract
BackgroundABO blood types have widespread clinical use and robust associations with cardiovascular disease. Many studies determine ABO blood types using tag single nucleotide polymorphisms (tSNPs) to characterize functional variation. However, tSNPs with low linkage disequilibrium (LD) may promote misinference of ABO blood types, particularly in diverse populations. MethodsBibliographic databases were searched for studies (2005-2022) using tSNPs to determine ABO alleles in accordance with PRISMA 2020 guidelines. We calculated linkage between tSNPs and functional variants across inferred continental ancestry groups from 1000 Genomes (AFR, AMR, EAS, EUR). We compared r2 across ancestry and assessed real-world consequences by comparing tSNP-derived blood types to serology in a large, diverse population from the All of Us Research Program (AoURP). ResultsWe observe a lack of phasing and frequent use of inappropriate tSNPs in blood type determination, particularly for O alleles. Linkage between functional variants and O allele tSNPs was significantly lower in African (median r2=0.443) compared to East Asian (r2=0.946, p=1.1x10-5) and European (r2=0.869, p=0.023). In AoURP, discordance between tSNP-derived blood types and serology was high across all SNPs in African ancestry individuals and linkage was strongly correlated with discordance across all ancestries ({rho}=-0.90, p=3.08x10-23). ConclusionWe observe common use of inappropriate tSNPs to determine ABO blood type, particularly for O alleles and with some tSNPs mistyping up to 58% of individuals. Our results highlight the lack of transferability of tSNPs across ancestries and potential exacerbation of disparities in genomic research for underrepresented populations.
Ljungdahl, A.; Kohani, S.; Page, N. F.; Wells, E. S.; Wigdor, E. M.; Dong, S.; Sanders, S. J.
Show abstract
Missense variants that alter a single amino acid in the encoded protein contribute to many human disorders but pose a substantial challenge in interpretation. Though these variants can be reliably identified through sequencing, distinguishing the clinically significant ones remains difficult, such that "Variants of Unknown Significance" outnumber those classified as "Pathogenic" or "Likely Pathogenic." Numerous in silico approaches have been developed to predict the functional impact of missense variants to inform clinical interpretation, the latest being AlphaMissense, which uses artificial intelligence methods trained on predicted protein structure. To independently assess the performance of AlphaMissense and 38 other predictors of missense severity, we compared predictions to data from multiplexed assays of variant effect (MAVE). MAVE experiments generate almost every possible individual amino acid change in a gene and measure their functional impact using a high-throughput assay. Assessing 17,696 variants across five genes (DDX3X, MSH2, PTEN, KCNQ4, and BRCA1), we find that AlphaMissense is consistently one of the top five algorithms based on correlation with functional impact and is the best-correlated algorithm for two genes. We conclude that AlphaMissense represents the current best-in-class predictor by this metric; however, the improvement over other algorithms is modest. We note that multiple missense predictors, including AlphaMissense, appear to overcall variants as pathogenic despite minimal functional impact and that substantially more high-quality training data, including consistently analyzed patient cohorts and MAVE analyses, are required to improve accuracy.
Dill-McFarland, K. A.; Andrade, B. B.; Figueiredo, M. C.; Andrade, A. M.; Avendano-Rangel, F.; Cordeiro-Santos, M.; Kritski, A. L.; Rolla, V. C.; Cubillos-Angulo, J. M.; Kalams, S. A.; Simmons, J. D.; Oakes, J. M.; Pena Avila, J.; Nakaya, H. I.; Gangula, R. D.; Rebeiro, P. F.; Amorim, G.; Mallal, S. A.; Regional Prospective Observational Research in Tuberculosis (RePORT)-Brazil Consortium, ; Sterling, T. R.; Hawn, T. R.
Show abstract
Although genetic factors contribute to tuberculosis (TB) risk, no cross-population causal variants have been identified by genome-wide association studies (GWAS). Here, we utilized low-pass whole genome sequencing (lpWGS) with imputation plus detailed epidemiologic risk factors and single-cell expression quantitative loci (sceQTL) to address prior GWAS limitations. Using 947 pulmonary tuberculosis (PTB) cases and 1807 close contact controls in the Regional Prospective Observational Research in TB (RePORT) study in Brazil, we estimated PTB heritability to be 47.7%. We identified 19 SNPs associated with PTB (P<5E-8) after adjustment for major risk factors (HIV, diabetes, smoking). Seven of these SNPs were associated with peripheral blood cell-specific sceQTLs in controls. Specifically, SNPs cis to transcription factors ZNF717 and MAML3 were associated with PTB disease and gene expression in monocytes, T cells, or B cells. Overall, this study utilized lpWGS, in-depth epidemiology, and single-cell analyses to detect population-specific genetic risk factors for PTB in Brazil. SUMMARYRobust correction for tuberculosis risk factors in GWAS in combination with paired single-cell transcriptomics reveals novel genetic risk of pulmonary tuberculosis with measurable consequences for baseline gene expression in multiple cell types.
Holla, B.; Mahadevan, J.; Ganesha, S.; Sud, R.; Janardhanan, M.; Balachander, S.; Strom, N.; Mattheisen, M.; Sullivan, P. F.; Huang, H.; Zandi, P.; Benegal, V.; Reddy, J. Y.; Jain, S.; CVEDA collaborators, ; ADBS-CBM consortium, ; iPSYCH OCD consortium, ; NORDiC OCD & Related Disorders Consortium, ; Purushottam, M.; Viswanath, B.
Show abstract
Genome-wide association studies across diverse populations may help validate and confirm genetic contributions to risk of disease. We estimated the extent of population stratification as well as the predictive accuracy of polygenic scores (PGS) derived from European samples to a data set from India. We analysed 2685 samples from two data sets, a population neurodevelopmental study (cVEDA) and a hospital-based sample of bipolar affective disorder (BD) and obsessive-compulsive disorder (OCD). Genotyping was conducted using Illuminas Global Screening Array. Population structure was examined with principal component analysis (PCA), uniform manifold approximation and projection (UMAP), support vector machine (SVM) ancestry predictions, and admixture analysis. PGS were calculated from the largest available European discovery GWAS summary statistics for BD, OCD, and externalizing traits using two Bayesian methods that incorporate local linkage disequilibrium structures (PGS-CS-auto) and functional genomic annotations (SBayesRC). Our analyses reveal global and continental PCA overlap with other South Asian populations. Admixture analysis revealed a north-south genetic axis within India (FST 1.6%). The UMAP partially reconstructed the contours of the Indian subcontinent. The Bayesian PGS analyses indicates moderate-to-high predictive power for BD. This was despite the cross-ancestry bias of the discovery GWAS dataset, with the currently available data. However, accuracy for OCD and externalizing traits was much lower. The predictive accuracy was perhaps influenced by the sample size of the discovery GWAS and phenotypic heterogeneity across the syndromes and traits studied. Our study results highlight the accuracy and generalizability of newer PGS models across ancestries. Further research, across diverse populations, would help understand causal mechanisms that contribute to psychiatric syndromes and traits.
Wang, X.; Sofer, T.; Frei, O.; Kaplan, R.; Perreira, K. M.; Franceschini, N.; Parada, H.; Zhou, L.; Andreassen, O. A.; Gonzalez, H.; Dale, A. M.; Broce, I. J.
Show abstract
Polygenic scores (PGS) offer moderate to high prediction accuracy for complex traits, but most are developed in European ancestry cohorts, reducing their performance in populations of other ancestries. This study aimed to improve standing height prediction, a heritable and ancestry-influenced trait, in an admixed Latino cohort (HCHS/SOL) by modeling ancestry using principal components (PCs) alongside PGS. SNPs were selected from a large European ancestry GWAS using various p-value thresholds, and weights were trained using traditional and penalized regression in the UK Biobank (UKB). PGS with PCs were trained separately in HCHS/SOL and UKB. Compared to PGS alone, modeling PGS with PCs substantially improved height prediction in HCHS/SOL (R{superscript 2} increase of [~]0.1), while mild improvements were observed in UKB (R{superscript 2} increase of [~]0.01). These results underscore the importance of incorporating genetic ancestry into predictive models for admixed populations, particularly when the trait exhibits ancestry-specific associations.
Grotzinger, A. D.; de la Fuente, J.; Davies, G.; Nivard, M. G.; Tucker-Drob, E. M.
Show abstract
Spearmans observation in 1904 that distinct cognitive functions--such as reasoning, processing speed, and episodic memory--are positively intercorrelated has given rise to over a century of speculation and investigation into their common and domain-specific mechanisms of variation. Here we develop and validate Transcriptome-wide Structural Equation Modeling (T-SEM), a novel method for studying the effects of tissue-specific gene expression within multivariate space. We apply T-SEM to investigate the shared and unique functional genomic characteristics of seven, distinct cognitive traits (N = 11,263-331,679). We identify 184 genes associated with general cognitive function (g), including 10 novel genes not identified in univariate analysis for the individual cognitive traits. We go on to apply Stratified Genomic SEM to identify enrichment for g within 29 functional genomic categories. This includes categories indexing the intersection of protein-truncating variant intolerant (PI) genes and specific neuronal cell types, which we also find to be enriched for the genetic covariance between g and a psychotic disorders factor.
Parrish, R. L.; Buchman, A. S.; Tasaki, S.; Wang, Y.; Avey, D.; Xu, J.; De Jager, P. L.; Bennett, D. A.; Epstein, M. P.; Yang, J.
Show abstract
Multiple reference panels of a given tissue or multiple tissues often exist, and multiple regression methods could be used for training gene expression imputation models for TWAS. To leverage expression imputation models (i.e., base models) trained with multiple reference panels, regression methods, and tissues, we develop a Stacked Regression based TWAS (SR-TWAS) tool which can obtain optimal linear combinations of base models for a given validation transcriptomic dataset. Both simulation and real studies showed that SR-TWAS improved power, due to increased effective training sample sizes and borrowed strength across multiple regression methods and tissues. Leveraging base models across multiple reference panels, tissues, and regression methods, our real application studies identified 6 independent significant risk genes for Alzheimers disease (AD) dementia for supplementary motor area tissue and 9 independent significant risk genes for Parkinsons disease (PD) for substantia nigra tissue. Relevant biological interpretations were found for these significant risk genes.
Ramdas, S.; Kahali, B.
Show abstract
The APOE {varepsilon}4 allele is the strongest genetic risk factor for Alzheimers Disease. However, its distribution across Indian populations is poorly characterized. We analyze APOE allele frequencies in 9,524 individuals from 83 distinct populations in the GenomeIndia dataset. {varepsilon}4 frequencies show large variation across populations within India, ranging from 2.7% to 36.1%, with a median of 11%. Tribal populations have higher {varepsilon}4 frequencies compared to non-tribal groups, while Tibeto-Burman populations have significantly lower frequencies. One tribal population from the northern coastal highlands has {varepsilon}4 frequency of 0.36, with 59% of individuals being carriers. {varepsilon}4 carrier status correlates significantly with lipid phenotypes including LDL, HDL, total cholesterol, and triglycerides. Collectively, these findings reveal exceptional genetic diversity in Alzheimers Disease risk across India and have important implications for population-specific screening strategies, genetic counseling, and precision medicine approaches to dementia prevention.
Zhao, Z.; Yi, Y.; Wu, Y.; Zhong, X.; Lin, Y.; Hohman, T. J.; Fletcher, J.; Lu, Q.
Show abstract
Polygenic risk scores (PRSs) have wide applications in human genetics research. Notably, most PRS models include tuning parameters which improve predictive performance when properly selected. However, existing model-tuning methods require individual-level genetic data as the training dataset or as a validation dataset independent from both training and testing samples. These data rarely exist in practice, creating a significant gap between PRS methodology and applications. Here, we introduce PUMAS (Parameter-tuning Using Marginal Association Statistics), a novel method to fine-tune PRS models using summary statistics from genome-wide association studies (GWASs). Through extensive simulations, external validations, and analysis of 65 traits, we demonstrate that PUMAS can perform a variety of model-tuning procedures (e.g. cross-validation) using GWAS summary statistics and can effectively benchmark and optimize PRS models under diverse genetic architecture. On average, PUMAS improves the predictive R2 by 205.6% and 62.5% compared to PRSs with arbitrary p-value cutoffs of 0.01 and 1, respectively. Applied to 211 neuroimaging traits and Alzheimers disease, we show that fine-tuned PRSs will significantly improve statistical power in downstream association analysis. We believe our method resolves a fundamental problem without a current solution and will greatly benefit genetic prediction applications.