Back

Human Genetics and Genomics Advances

Elsevier BV

All preprints, ranked by how well they match Human Genetics and Genomics Advances's content profile, based on 84 papers previously published here. The average preprint has a 0.08% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Improving imputation quality in Samoans through the integration of population-specific sequences into existing reference panels

Carlson, J. C.; Krishnan, M.; Liu, S.; Anderson, K. J.; Zhang, J. Z.; Yapp, T.-A. J.; Chiyka, E. A.; Dikec, D. A.; Cheng, H.; Naseri, T.; Reupena, M. S.; Viali, S.; Deka, R.; Hawley, N. L.; McGarvey, S. T.; Weeks, D. E.; Minster, R. L.

2023-10-31 genetic and genomic medicine 10.1101/2023.10.31.23297835 medRxiv
Top 0.1%
38.4%
Show abstract

Genotype imputation is fundamental to association studies, and yet even gold standard panels like TOPMed are limited in the populations for which they yield good imputation. Specifically, Pacific Islanders are poorly represented in extant panels. To address this, we constructed an imputation reference panel using 1,285 Samoan individuals with whole-genome sequencing, combined with 1000 Genomes Project (1KGP) individuals, to create a reference panel that better represents Pacific Islander, specifically Samoan, genetic variation. We compared this panel to 1KGP and TOPMed-R3 panels based on imputed variants using genotyping array data for 1,834 Samoan participants who were not part of the panels. The 1KGP + 1285 Samoan panel yielded up to two times more well-imputed (r2 [≥] 0.80) variants than TOPMed-R3 and 1KGP and was enriched for moderate and high impact variants. There was improved imputation accuracy across the minor allele frequency (MAF) spectrum, although it was most pronounced for variants with 0.01 [≤] MAF [≤] 0.05. Imputation accuracy (r2) was greater for population-specific variants (high fixation index, FST) and those from larger haplotypes (high LD score). However, the gain in imputation accuracy over TOPMed-R3 was largest for small haplotypes (low LD score), reflecting the Samoan panels ability to capture population-specific variation not well tagged by other panels. We also augmented the 1KGP reference panel with varying numbers of Samoan participants and found that panels with 24 Samoans yielded similar performance to TOPMed-R3, and panels with 48 or more Samoans included outperformed TOPMed-R3 for all variants with MAF [≥] 0.001. Meta imputation of the TOPMed-R3 and 1285 Samoan panels yielded poorer performance than the Samoan only panel. We also demonstrated that the phasing of the reference panel impacts the imputation of population-specific variants when the reference panel is composed of individuals from an isolated population and not combined with ancestrally diverse haplotypes. This study identifies variants with improved imputation using population-specific reference panels and provides a framework for constructing other population-specific reference panels.

2
Genome admixture analysis of 1,030 Ugandan infants with neonatal sepsis and hydrocephalus demonstrates geographical stratification of population disease risk

Movassagh, M.; Newbury, L.; Hehnly, C.; Whalen, A.; Peterson, M.; Mondragon Estrada, E.; Ericson, J.; Smith, J.; Sasanami, M.; Natukwatsa, D.; Mugamba, J.; Ssenyonga, P.; Onen, J.; Burgoine, K.; Zhang, L.; Olupot-Olupot, P.; Kumbakumba, E.; Wegoye, E.; Ochora, M.; Mulondo, R.; Mbabazi-Kabachelor, E.; Fronterre, C.; Broach, J.; Paulson, J.; Morton, S.; Schiff, S.

2026-03-23 pediatrics 10.64898/2026.03.16.26348489 medRxiv
Top 0.1%
38.1%
Show abstract

BackgroundNeonatal disorders such as post-infectious hydrocephalus exhibit a higher incidence in Africa, where the intricate relationships between genetic ancestry, environmental exposures, and other risk factors likely contribute to the increased incidence. MethodsTo start to characterize the common genetic architecture of Ugandan infants, we analyzed genome sequencing data from 1,030 Ugandan infants recruited from studies targeting neonatal sepsis and hydrocephalus. We employed genetic admixture analysis and integrated geospatial data to examine the relationships between genetic backgrounds and disease prevalence within this cohort. ResultsOur results identified four distinct genetic admixture groups, each correlating strongly with specific geographic distributions across Uganda. Notably, a predominance of one admixture group, most common in northern Uganda, was overrepresented in the participants with post-infectious hydrocephalus. ConclusionThis study underscores the importance of genetic factors in disease manifestation at the population level, and a role for such precision public health approaches in complex neonatal disorders in African populations.

3
Evaluating Polygenic Score Transferability for Lipid Traits in Underrepresented Populations: Evidence from Samoan Cohorts

Yapp, T.-A. J.; Krishnan, M.; Liu, S.; Manna, S. L.; Cheng, H.; Naseri, T.; Reupena, M. S.; Viali, S.; Tuitele, J.; Deka, R.; Hawley, N. L.; McGarvey, S. T.; Weeks, D. E.; Minster, R. L.; Carlson, J. C.

2026-06-30 genetic and genomic medicine 10.64898/2026.06.26.26356725 medRxiv
Top 0.1%
34.9%
Show abstract

Dyslipidemia is a significant risk factor for cardiovascular disease (CVD), the leading cause of death in Samoa, accounting for 34% of deaths. Polygenic scores (PGS) derived from large scale multi ancestry genome-wide association studies offer potential for improved CVD risk prediction by aggregating genetic effects on lipid traits, yet their performance in Pacific Islander populations remains largely unknown. We evaluated the transferability of multi-ancestry PGS for LDL cholesterol (LDL C), HDL cholesterol (HDL C), triglycerides (TG), and total cholesterol (TC) in 4,342 Samoan adults across five cohorts spanning 1990 to 2010. PGS derived from Graham et al. and Kanoni et al. multi-ancestry meta-analyses were harmonized with genome-wide imputed genotypes using a Samoan-specific reference panel, and performance was assessed using incremental R^2 from linear mixed models with bootstrapped confidence intervals. PGS performance varied across traits and cohorts: HDL C showed the highest performance (incremental R^2 5.0 to15.0%), followed by LDL C (5.7 to 8.6%) and TC (5.0 to10.7%), with TG showing the lowest performance (3.5 to 7.0%). Meaningful LDL C transferability was achieved only when using a genome-wide PRS CS score (99.6 to 99.7% variant matching), whereas a curated pruning-and-thresholding score achieved only ~9% matching and near-zero performance. These findings establish the first systematic benchmarks for lipid PGS performance in Samoans, demonstrate that multi-ancestry scores can achieve meaningful transferability in this underrepresented population when genome-wide variant coverage is ensured, and highlight the importance of rigorous variant harmonization assessment prior to clinical deployment of PGS in diverse populations.

4
A Novel Method to Disentangle Tightly Linked Risk and Resilience Genes for Brain Disorders Application to Alzheimers Disease

Barnett, E. J.; Hess, J. L.; Hou, J.; Escott-Price, V.; Fennema-Notestine, C.; Kremen, W.; Lin, S.-J.; Zhang, C.; Gaiteri, C.; Elman, J.; Holmans, P.; Faraone, S. V.; Glatt, S. J.

2025-02-27 genetic and genomic medicine 10.1101/2025.02.26.25322962 medRxiv
Top 0.1%
34.2%
Show abstract

Genetic risk factors for neuropsychiatric disorders are well documented. However, some individuals with high genetic risk remain unaffected, and the mechanisms underlying such resilience remain poorly understood. The presence of protective resilience factors that mitigate risk could help explain the disconnect between predicted risk and reality, particularly when genetic contributions are substantial but incompletely understood. Identifying and studying resilience factors could improve our understanding of pathology, enhance risk prediction, and inform preventive measures or treatment strategies. However, such efforts are complicated by the difficulty of identifying resilience that is separable from low risk. We developed a novel adversarial multi-task neural network model to detect genetic resilience markers. The model learns to separate high-risk unaffected individuals from affected individuals at similar risk while "unlearning" patterns found in low-risk groups using adversarial learning. In simulated and existing Alzheimers disease (AD) datasets, we identified markers of resilience with a feature-importance-based approach that prioritized specificity, generated resilience scores, and analyzed associations with polygenic risk scores (PRS). In simulations, our model had high specificity and sensitivity in identifying resilience markers, significantly outperforming traditional approaches. Applied to AD data, the model generated genetic resilience scores protective against AD and independent of PRS. We identified five resilience-associated SNPs, including known AD-associated variants, underscoring their potential involvement in resilience. Our findings support the utility of resilience scores in modifying risk predictions, particularly for high-risk groups. Expanding this method could aid in understanding resilience mechanisms, potentially improving diagnosis, prevention, and treatment strategies for AD and other brain disorders.

5
The accuracy of polygenic score models for anthropometric traits and Type II Diabetes in the Native Hawaiian Population

Lo, Y.-C.; Chan, T. F.; Jeon, S.; Maskarinec, G.; Taparra, K.; Nakatsuka, N.; Yu, M.; Chen, C.-Y.; Lin, Y.-F.; Wilkens, L. R.; Le Marchand, L.; Haiman, C. A.; Chiang, C. W. K.

2023-12-28 genetic and genomic medicine 10.1101/2023.12.25.23300499 medRxiv
Top 0.1%
33.7%
Show abstract

Polygenic scores (PGS) are promising in stratifying individuals based on the genetic susceptibility to complex diseases or traits. However, the accuracy of PGS models, typically trained in European- or East Asian-ancestry populations, tend to perform poorly in other ethnic minority populations, and their accuracies have not been evaluated for Native Hawaiians. Using body mass index, height, and type-2 diabetes as examples of highly polygenic traits, we evaluated the prediction accuracies of PGS models in a large Native Hawaiian sample from the Multiethnic Cohort with up to 5,300 individuals. We evaluated both publicly available PGS models or genome-wide PGS models trained in this study using the largest available GWAS. We found evidence of lowered prediction accuracies for the PGS models in some cases, particularly for height. We also found that using the Native Hawaiian samples as an optimization cohort during training did not consistently improve PGS performance. Moreover, even the best performing PGS models among Native Hawaiians would have lowered prediction accuracy among the subset of individuals most enriched with Polynesian ancestry. Our findings indicate that factors such as admixture histories, sample size and diversity in GWAS can influence PGS performance for complex traits among Native Hawaiian samples. This study provides an initial survey of PGS performance among Native Hawaiians and exposes the current gaps and challenges associated with improving polygenic prediction models for underrepresented minority populations.

6
Modeling the longitudinal changes of ancestry diversity in the Million Veteran Program

Wendt, F. R.; Pathak, G. A.; Vahey, J.; Qin, X.; Koller, D.; Cabrera-Mendoza, B.; Haeny, A.; Harrington, K. M.; Rajeevan, N.; Duong, L. M.; Levey, D. F.; De Angelis, F.; De Lillo, A.; Bigdeli, T. B.; Pyarajan, S.; VA Million Veteran Program, ; Gaziano, J. M.; Gelernter, J.; Aslan, M.; Provenzale, D.; Helmer, D. A.; Hauser, E. R.; Polimanti, R.; Department of Veteran Affairs Cooperative Study Program (#2006),

2022-01-25 genomics 10.1101/2022.01.24.477583 medRxiv
Top 0.1%
33.6%
Show abstract

The Million Veteran Program (MVP) participants represent 100 years of US history, including significant social and demographic change over time. Our study assessed two aspects of the MVP: (i) longitudinal changes in population diversity and (ii) how these changes can be accounted for in genome-wide association studies (GWAS). The MVP was divided into five birth cohorts (N-range=123,888 [born from 1943-1947] to 136,699 [born from 1948-1953]). Groups of participants were defined by (i) HARE (harmonized ancestry and race/ethnicity) and (ii) a random-forest clustering approach using the 1000 Genomes Project and the Human Genome Diversity Project (1kGP+HGDP) reference panels (77 world populations representing six continental groups). In these groups, we performed GWASs of height, a trait potentially affected by population stratification. Birth cohorts demonstrate important trends in ancestry diversity over time. More recent HARE-assigned Europeans, Africans, and Hispanics had lower European ancestry proportions than older birth cohorts (0.010<Cohens d<0.259, p<7.80x10-4). Conversely, HARE-assigned East Asians showed an increase in European ancestry proportion over time. In GWAS of height using HARE assignments, genomic inflation due to population stratification was prevalent across all birth cohorts (linkage disequilibrium score regression intercept=1.08{+/-}0.042). The 1kGP+HGDP-based ancestry assignment significantly reduced the population stratification (mean intercept reduction=0.045{+/-}0.007, p<0.05) confounding in the GWAS statistics. This study provides a comprehensive characterization of ancestry diversity of the MVP cohort over time and highlights that more refined modeling of genetic diversity (e.g., the 1kGP+HGDP-based ancestry assignment) can more accurately capture the polygenic architecture of traits and diseases that could be affected by population stratification.

7
Structural Variant Imputation in Samoans Using a Population-Specific Reference Panel

Spor, L. M.; Liau, E. M.; Sanchis-Juan, A.; Silva, A. N.; Cheng, H.; Naseri, T.; Reupena, M. S.; Viali, S.; Tuitele, J.; Kershaw, E. E.; Deka, R. D.; Hawley, N. L.; McGarvey, S. T.; Weeks, D. E.; Carlson, J. C.; Brand, H.; Minster, R. L.

2026-07-10 genetic and genomic medicine 10.64898/2026.07.07.26357461 medRxiv
Top 0.1%
31.7%
Show abstract

Structural variants (SVs) are often excluded from genetic research because they are difficult to call, but they can have substantial effects on phenotypic traits. SVs have not previously been characterized in Samoans, an understudied population with a high burden of complex diseases. Using short-read whole genome sequencing data, we called SVs in 1,276 Samoans and created a Samoan-specific imputation panel inclusive of both SVs and single nucleotide variants (SNVs), called the Soifua Manuia-SV panel. Using this panel, we imputed SVs and SNVs in 3,611 Samoans with array data, enabling analysis of SV-phenotype associations in a sample of 4,887 Samoan participants. We evaluated imputation performance in Samoans against two other reference panels: (i) an SNV-only Samoan-specific reference panel, to assess whether SV inclusion impacts SNV imputation, and (ii) an SV and SNV, multi-ancestry reference panel composed of 1000 Genomes participants, which did not include Polynesians, to assess the importance of including the target population in the reference panel. The Soifua Manuia-SV panel substantially outperformed the multi-ancestry SV and SNV panel, yielding 5.5 million more high-quality (r2[&ge;]0.8) variants, including over 8,000 more high-quality SVs. SNV imputation based on the two Samoan-specific panels performed similarly overall, suggesting that SV inclusion does not strongly impact SNV imputation quality. This work highlights the importance of population representation for accurate imputation.

8
The Polygenic Score Catalog: new functionality and tools to enable FAIR research

Lambert, S. A.; Wingfield, B.; Gibson, J. T.; Gil, L.; Ramachandran, S.; Yvon, F.; Saverimuttu, S.; Tinsley, E.; Lewis, E.; Ritchie, S. C.; Wu, J.; Canovas, R.; McMahon, A.; Harris, L. W.; Parkinson, H.; Inouye, M.

2024-05-31 genetic and genomic medicine 10.1101/2024.05.29.24307783 medRxiv
Top 0.1%
31.6%
Show abstract

Polygenic scores (PGS) have transformed human genetic research and have multiple potential clinical applications, including risk stratification for disease prevention and prediction of treatment response. Here, we present a series of recent enhancements to the PGS Catalog (www.PGSCatalog.org), the largest findable, accessible, interoperable, and reusable (FAIR) repository of PGS. These include expansions in data content and ancestral diversity as well as the addition of new features. We further present the PGS Catalog Calculator (pgsc_calc, https://github.com/PGScatalog/pgsc_calc), an open-source, scalable and portable pipeline to reproducibly calculate PGS that securely democratizes equitable PGS applications by implementing genetic ancestry estimation and score normalization using reference data. With the PGS Catalog & calculator users can now quantify an individuals genetic predisposition for hundreds of common diseases and clinically relevant traits. Taken together, these updates and tools facilitate the next generation of PGS, thus lowering barriers to the clinical studies necessary to identify where PGS may be integrated into clinical practice.

9
Genetic and demographic predictors of general reading ability in two cohorts

Lancaster, H. S.; Dinu, V.; Li, J.; Gruen, J. R.; GRaD Consortium,

2021-08-29 pediatrics 10.1101/2021.08.24.21262573 medRxiv
Top 0.1%
31.4%
Show abstract

PurposeReading ability is a complex skill utilizing multiple proficiencies and that develops through interactions between genetic and environmental factors. This study presents an alternative analytic pipeline to identify key genetic and demographic contributors to reading ability. MethodsWe analyzed data from the Avon Longitudinal Study of Parents and Children (ALSPAC; N = 3 232) using a multi-step analytical pipeline. To reduce measurement error, we generated a latent reading ability score. We selected single nucleotide polymorphisms (SNPs) based on existing literature and genome-wide association studies (GWAS). We applied elastic net regression to identify informative predictors in two models, a SNP-only model and a SNP-plus demographic, environmental, and behavioral variables model. We compared the SNP-based heritability estimates and R2 from the fitted models. We also performed pathway enrichment analysis on the informative SNPs. ResultsThe traditional GWAS identified one genome-wide significant SNP on chromosome X and produced a moderate heritability estimate of .23 (SE = 0.07). We included 148 SNPs in the elastic net models. The SNP-only model identified 61 informative SNPs (R2 = .12), whereas the SNP-plus model identified 96 informative SNPs (R2 = .32). The SNP-plus model also showed that several behavioral characteristics positively predicted latent reading ability. Enrichment analysis revealed overrepresentation of several biological pathways among the informative SNPs. ConclusionsThis study shows that our analytic pipeline can identify important genetic and demographic predictors of reading ability, providing a powerful alternative to traditional methods and contributing to a deeper understanding of the factors that drive reading development.

10
Improving GWAS performance in underrepresented groups by appropriate modeling of genetics, environment, and sociocultural factors

Cataldo-Ramirez, C.; Lin, M.; McMahon, A.; Gignoux, C.; Weaver, T. D.; Henn, B. M.

2026-04-08 genetics 10.1101/2024.10.28.620716 medRxiv
Top 0.1%
30.6%
Show abstract

Genome-wide association studies (GWAS) and polygenic score (PGS) development are typically constrained by the data available in biobank repositories in which European cohorts are vastly overrepresented. Here, we increase the utility of non-European participant data within the UK Biobank (UKB) by characterizing the genetic affinities of UKB participants who self-identify as Bangladeshi, Indian, Pakistani, "White and Asian" (WA), and "Any Other Asian" (AOA), towards creating a more robust South Asian sample size for future genetic analyses. We assess the relationships between genetic structure and self-selected ethnic identities and use consistent patterns of clustering in the dataset to train a support vector machine (SVM). The SVM was utilized to reassign n = 1,853 AOA and WA participants at the subcontinental level, and increase the sample size of the UKB South Asian group by 1,381 additional participants. We further leverage these samples to assess GWAS performance and PGS development. We include environmental covariates in the height GWAS by implementing a rigorous covariate selection procedure, and compare the outputs of two GWAS models: GWASnull and GWASenv. We show that PGS performance derived from both GWAS models yield comparable prediction to PGS models developed with an order of magnitude larger training, and environmentally-adjusted PGS models reduce the sex-bias in predictive performance. In summary, we demonstrate how GWAS performance can be improved by leveraging ambiguous ethnicity codes, ancestry matched imputation panels, and including environmental covariates.

11
Within- and Between-Family Validation of Nine Polygenic Risk Scores Developed in 1.5 Million Individuals: Implications for IVF, Embryo Selection, and Reduction in Lifetime Disease Risk

Cordogan, S.; Starr, D. B.; Treff, N. R.; Lanchbury, J.; Goldstein, E.; Sadeghi, K.; Burmeister, M.; Dayani, L.; Macdonald, P.; Keen-Kim, J. D.; Fishel, S.; Cervantes, E.; Folkersen, L.

2025-10-28 genetic and genomic medicine 10.1101/2025.10.24.25338613 medRxiv
Top 0.1%
30.5%
Show abstract

Polygenic risk scores (PRSs) can reduce lifetime disease risk by guiding embryo selection during in vitro fertilization (IVF). We performed genome-wide association meta-analyses totaling [~]1.5 million individuals to construct state-of-the-art PRSs for nine diseases: Alzheimers disease, breast cancer, coronary artery disease, endometriosis, hypertension, prostate cancer, rheumatoid arthritis, type 1 diabetes, and type 2 diabetes. The resulting predictors achieved liability-scale R2 values of up to 22.9% for type 2 diabetes, matching or exceeding previously published benchmarks across all scores. Three PRS - Alzheimers disease, prostate cancer, and type 2 diabetes - explained over 75% of the common-SNP heritability. Within-family validation in 40,872 siblings across 18,840 families showed that, for eight of nine diseases, predictive performance was comparable to population-level results, confirming substantial direct genetic effects. Modeling of embryo selection suggests that couples with five euploid embryos could achieve 27-67% relative risk reduction across diseases. While the limitations of genetic data availability meant that these estimates were performed in European ancestry samples, we performed validation in the multi-ancestry US-based All of Us Biobank, demonstrating significant statistical power across ancestries. These findings support the clinical applicability of PRS-guided embryo selection to reduce the burden of common diseases.

12
Disparities in ABO Blood Type Determination Across Diverse Ancestries: A Systematic Review and Validation in the All of Us Research Program

Martinez, K. L.; Klein, A.; Martin, J. R.; Sampson, C. U.; Giles, J. B.; Beck, M. L.; Bhakta, K.; Quatraro, G.; Farol, J.; Karnes, J. H.

2024-02-24 genetic and genomic medicine 10.1101/2024.02.21.24302372 medRxiv
Top 0.1%
27.9%
Show abstract

BackgroundABO blood types have widespread clinical use and robust associations with cardiovascular disease. Many studies determine ABO blood types using tag single nucleotide polymorphisms (tSNPs) to characterize functional variation. However, tSNPs with low linkage disequilibrium (LD) may promote misinference of ABO blood types, particularly in diverse populations. MethodsBibliographic databases were searched for studies (2005-2022) using tSNPs to determine ABO alleles in accordance with PRISMA 2020 guidelines. We calculated linkage between tSNPs and functional variants across inferred continental ancestry groups from 1000 Genomes (AFR, AMR, EAS, EUR). We compared r2 across ancestry and assessed real-world consequences by comparing tSNP-derived blood types to serology in a large, diverse population from the All of Us Research Program (AoURP). ResultsWe observe a lack of phasing and frequent use of inappropriate tSNPs in blood type determination, particularly for O alleles. Linkage between functional variants and O allele tSNPs was significantly lower in African (median r2=0.443) compared to East Asian (r2=0.946, p=1.1x10-5) and European (r2=0.869, p=0.023). In AoURP, discordance between tSNP-derived blood types and serology was high across all SNPs in African ancestry individuals and linkage was strongly correlated with discordance across all ancestries ({rho}=-0.90, p=3.08x10-23). ConclusionWe observe common use of inappropriate tSNPs to determine ABO blood type, particularly for O alleles and with some tSNPs mistyping up to 58% of individuals. Our results highlight the lack of transferability of tSNPs across ancestries and potential exacerbation of disparities in genomic research for underrepresented populations.

13
AlphaMissense is better correlated with functional assays of missense impact than earlier prediction algorithms

Ljungdahl, A.; Kohani, S.; Page, N. F.; Wells, E. S.; Wigdor, E. M.; Dong, S.; Sanders, S. J.

2023-10-27 genetics 10.1101/2023.10.24.562294 medRxiv
Top 0.1%
27.0%
Show abstract

Missense variants that alter a single amino acid in the encoded protein contribute to many human disorders but pose a substantial challenge in interpretation. Though these variants can be reliably identified through sequencing, distinguishing the clinically significant ones remains difficult, such that "Variants of Unknown Significance" outnumber those classified as "Pathogenic" or "Likely Pathogenic." Numerous in silico approaches have been developed to predict the functional impact of missense variants to inform clinical interpretation, the latest being AlphaMissense, which uses artificial intelligence methods trained on predicted protein structure. To independently assess the performance of AlphaMissense and 38 other predictors of missense severity, we compared predictions to data from multiplexed assays of variant effect (MAVE). MAVE experiments generate almost every possible individual amino acid change in a gene and measure their functional impact using a high-throughput assay. Assessing 17,696 variants across five genes (DDX3X, MSH2, PTEN, KCNQ4, and BRCA1), we find that AlphaMissense is consistently one of the top five algorithms based on correlation with functional impact and is the best-correlated algorithm for two genes. We conclude that AlphaMissense represents the current best-in-class predictor by this metric; however, the improvement over other algorithms is modest. We note that multiple missense predictors, including AlphaMissense, appear to overcall variants as pathogenic despite minimal functional impact and that substantially more high-quality training data, including consistently analyzed patient cohorts and MAVE analyses, are required to improve accuracy.

14
Genome-wide association study in Brazil identifies risk factor-adjusted genetic susceptibility to pulmonary tuberculosis with cell-specific gene expression effects

Dill-McFarland, K. A.; Andrade, B. B.; Figueiredo, M. C.; Andrade, A. M.; Avendano-Rangel, F.; Cordeiro-Santos, M.; Kritski, A. L.; Rolla, V. C.; Cubillos-Angulo, J. M.; Kalams, S. A.; Simmons, J. D.; Oakes, J. M.; Pena Avila, J.; Nakaya, H. I.; Gangula, R. D.; Rebeiro, P. F.; Amorim, G.; Mallal, S. A.; Regional Prospective Observational Research in Tuberculosis (RePORT)-Brazil Consortium, ; Sterling, T. R.; Hawn, T. R.

2025-03-14 genetic and genomic medicine 10.1101/2025.03.13.25323932 medRxiv
Top 0.1%
26.5%
Show abstract

Although genetic factors contribute to tuberculosis (TB) risk, no cross-population causal variants have been identified by genome-wide association studies (GWAS). Here, we utilized low-pass whole genome sequencing (lpWGS) with imputation plus detailed epidemiologic risk factors and single-cell expression quantitative loci (sceQTL) to address prior GWAS limitations. Using 947 pulmonary tuberculosis (PTB) cases and 1807 close contact controls in the Regional Prospective Observational Research in TB (RePORT) study in Brazil, we estimated PTB heritability to be 47.7%. We identified 19 SNPs associated with PTB (P<5E-8) after adjustment for major risk factors (HIV, diabetes, smoking). Seven of these SNPs were associated with peripheral blood cell-specific sceQTLs in controls. Specifically, SNPs cis to transcription factors ZNF717 and MAML3 were associated with PTB disease and gene expression in monocytes, T cells, or B cells. Overall, this study utilized lpWGS, in-depth epidemiology, and single-cell analyses to detect population-specific genetic risk factors for PTB in Brazil. SUMMARYRobust correction for tuberculosis risk factors in GWAS in combination with paired single-cell transcriptomics reveals novel genetic risk of pulmonary tuberculosis with measurable consequences for baseline gene expression in multiple cell types.

15
A cross ancestry genetic study of psychiatric disorders from India

Holla, B.; Mahadevan, J.; Ganesha, S.; Sud, R.; Janardhanan, M.; Balachander, S.; Strom, N.; Mattheisen, M.; Sullivan, P. F.; Huang, H.; Zandi, P.; Benegal, V.; Reddy, J. Y.; Jain, S.; CVEDA collaborators, ; ADBS-CBM consortium, ; iPSYCH OCD consortium, ; NORDiC OCD & Related Disorders Consortium, ; Purushottam, M.; Viswanath, B.

2024-04-27 genetic and genomic medicine 10.1101/2024.04.25.24306377 medRxiv
Top 0.1%
26.4%
Show abstract

Genome-wide association studies across diverse populations may help validate and confirm genetic contributions to risk of disease. We estimated the extent of population stratification as well as the predictive accuracy of polygenic scores (PGS) derived from European samples to a data set from India. We analysed 2685 samples from two data sets, a population neurodevelopmental study (cVEDA) and a hospital-based sample of bipolar affective disorder (BD) and obsessive-compulsive disorder (OCD). Genotyping was conducted using Illuminas Global Screening Array. Population structure was examined with principal component analysis (PCA), uniform manifold approximation and projection (UMAP), support vector machine (SVM) ancestry predictions, and admixture analysis. PGS were calculated from the largest available European discovery GWAS summary statistics for BD, OCD, and externalizing traits using two Bayesian methods that incorporate local linkage disequilibrium structures (PGS-CS-auto) and functional genomic annotations (SBayesRC). Our analyses reveal global and continental PCA overlap with other South Asian populations. Admixture analysis revealed a north-south genetic axis within India (FST 1.6%). The UMAP partially reconstructed the contours of the Indian subcontinent. The Bayesian PGS analyses indicates moderate-to-high predictive power for BD. This was despite the cross-ancestry bias of the discovery GWAS dataset, with the currently available data. However, accuracy for OCD and externalizing traits was much lower. The predictive accuracy was perhaps influenced by the sample size of the discovery GWAS and phenotypic heterogeneity across the syndromes and traits studied. Our study results highlight the accuracy and generalizability of newer PGS models across ancestries. Further research, across diverse populations, would help understand causal mechanisms that contribute to psychiatric syndromes and traits.

16
Explicitly modeling genetic ancestry to improve polygenic prediction accuracy for height in a large, admixed cohort of US Latinos: Findings from HCHS/SOL

Wang, X.; Sofer, T.; Frei, O.; Kaplan, R.; Perreira, K. M.; Franceschini, N.; Parada, H.; Zhou, L.; Andreassen, O. A.; Gonzalez, H.; Dale, A. M.; Broce, I. J.

2025-03-23 genetic and genomic medicine 10.1101/2025.03.21.25324423 medRxiv
Top 0.1%
26.2%
Show abstract

Polygenic scores (PGS) offer moderate to high prediction accuracy for complex traits, but most are developed in European ancestry cohorts, reducing their performance in populations of other ancestries. This study aimed to improve standing height prediction, a heritable and ancestry-influenced trait, in an admixed Latino cohort (HCHS/SOL) by modeling ancestry using principal components (PCs) alongside PGS. SNPs were selected from a large European ancestry GWAS using various p-value thresholds, and weights were trained using traditional and penalized regression in the UK Biobank (UKB). PGS with PCs were trained separately in HCHS/SOL and UKB. Compared to PGS alone, modeling PGS with PCs substantially improved height prediction in HCHS/SOL (R{superscript 2} increase of [~]0.1), while mild improvements were observed in UKB (R{superscript 2} increase of [~]0.01). These results underscore the importance of incorporating genetic ancestry into predictive models for admixed populations, particularly when the trait exhibits ancestry-specific associations.

17
Transcriptome-wide and Stratified Genomic Structural Equation Modeling Identify Neurobiological Pathways Underlying General and Specific Cognitive Functions

Grotzinger, A. D.; de la Fuente, J.; Davies, G.; Nivard, M. G.; Tucker-Drob, E. M.

2021-05-04 genetic and genomic medicine 10.1101/2021.04.30.21256409 medRxiv
Top 0.1%
26.0%
Show abstract

Spearmans observation in 1904 that distinct cognitive functions--such as reasoning, processing speed, and episodic memory--are positively intercorrelated has given rise to over a century of speculation and investigation into their common and domain-specific mechanisms of variation. Here we develop and validate Transcriptome-wide Structural Equation Modeling (T-SEM), a novel method for studying the effects of tissue-specific gene expression within multivariate space. We apply T-SEM to investigate the shared and unique functional genomic characteristics of seven, distinct cognitive traits (N = 11,263-331,679). We identify 184 genes associated with general cognitive function (g), including 10 novel genes not identified in univariate analysis for the individual cognitive traits. We go on to apply Stratified Genomic SEM to identify enrichment for g within 29 functional genomic categories. This includes categories indexing the intersection of protein-truncating variant intolerant (PI) genes and specific neuronal cell types, which we also find to be enriched for the genetic covariance between g and a psychotic disorders factor.

18
SR-TWAS: Leveraging Multiple Reference Panels to Improve TWAS Power by Ensemble Machine Learning

Parrish, R. L.; Buchman, A. S.; Tasaki, S.; Wang, Y.; Avey, D.; Xu, J.; De Jager, P. L.; Bennett, D. A.; Epstein, M. P.; Yang, J.

2023-06-27 genetic and genomic medicine 10.1101/2023.06.20.23291605 medRxiv
Top 0.1%
26.0%
Show abstract

Multiple reference panels of a given tissue or multiple tissues often exist, and multiple regression methods could be used for training gene expression imputation models for TWAS. To leverage expression imputation models (i.e., base models) trained with multiple reference panels, regression methods, and tissues, we develop a Stacked Regression based TWAS (SR-TWAS) tool which can obtain optimal linear combinations of base models for a given validation transcriptomic dataset. Both simulation and real studies showed that SR-TWAS improved power, due to increased effective training sample sizes and borrowed strength across multiple regression methods and tissues. Leveraging base models across multiple reference panels, tissues, and regression methods, our real application studies identified 6 independent significant risk genes for Alzheimers disease (AD) dementia for supplementary motor area tissue and 9 independent significant risk genes for Parkinsons disease (PD) for substantia nigra tissue. Relevant biological interpretations were found for these significant risk genes.

19
APOE4 Allele Frequencies Show Dramatic Variation Across Indian Populations

Ramdas, S.; Kahali, B.

2026-04-13 genetic and genomic medicine 10.64898/2026.04.09.26350483 medRxiv
Top 0.1%
23.1%
Show abstract

The APOE {varepsilon}4 allele is the strongest genetic risk factor for Alzheimers Disease. However, its distribution across Indian populations is poorly characterized. We analyze APOE allele frequencies in 9,524 individuals from 83 distinct populations in the GenomeIndia dataset. {varepsilon}4 frequencies show large variation across populations within India, ranging from 2.7% to 36.1%, with a median of 11%. Tribal populations have higher {varepsilon}4 frequencies compared to non-tribal groups, while Tibeto-Burman populations have significantly lower frequencies. One tribal population from the northern coastal highlands has {varepsilon}4 frequency of 0.36, with 59% of individuals being carriers. {varepsilon}4 carrier status correlates significantly with lipid phenotypes including LDL, HDL, total cholesterol, and triglycerides. Collectively, these findings reveal exceptional genetic diversity in Alzheimers Disease risk across India and have important implications for population-specific screening strategies, genetic counseling, and precision medicine approaches to dementia prevention.

20
Fine-tuning Polygenic Risk Scores with GWAS Summary Statistics

Zhao, Z.; Yi, Y.; Wu, Y.; Zhong, X.; Lin, Y.; Hohman, T. J.; Fletcher, J.; Lu, Q.

2019-10-18 genetics 10.1101/810713 medRxiv
Top 0.1%
22.8%
Show abstract

Polygenic risk scores (PRSs) have wide applications in human genetics research. Notably, most PRS models include tuning parameters which improve predictive performance when properly selected. However, existing model-tuning methods require individual-level genetic data as the training dataset or as a validation dataset independent from both training and testing samples. These data rarely exist in practice, creating a significant gap between PRS methodology and applications. Here, we introduce PUMAS (Parameter-tuning Using Marginal Association Statistics), a novel method to fine-tune PRS models using summary statistics from genome-wide association studies (GWASs). Through extensive simulations, external validations, and analysis of 65 traits, we demonstrate that PUMAS can perform a variety of model-tuning procedures (e.g. cross-validation) using GWAS summary statistics and can effectively benchmark and optimize PRS models under diverse genetic architecture. On average, PUMAS improves the predictive R2 by 205.6% and 62.5% compared to PRSs with arbitrary p-value cutoffs of 0.01 and 1, respectively. Applied to 211 neuroimaging traits and Alzheimers disease, we show that fine-tuned PRSs will significantly improve statistical power in downstream association analysis. We believe our method resolves a fundamental problem without a current solution and will greatly benefit genetic prediction applications.