Back

The American Journal of Human Genetics

Elsevier BV

All preprints, ranked by how well they match The American Journal of Human Genetics's content profile, based on 234 papers previously published here. The average preprint has a 0.19% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Pathogenic variants in TMEM184B cause a neurodevelopmental syndrome via alteration of metabolic signaling

Chapman, K. A.; Ullah, F.; Yahiku, Z. A.; Kodiparthi, S. V.; Kellaris, G.; Correia, S. P.; Stodberg, T.; Sofokleous, C.; Marinakis, N.; Fryssira, H.; Tsoutsou, E.; Traeger-Synodinos, J.; Accogli, A.; Salpietro, V.; Striano, P.; Berger, S. I.; Pond, K. W.; Sirimulla, S.; Davis, E. E.; Bhattacharya, M. R.

2024-07-01 genetic and genomic medicine 10.1101/2024.06.27.24309417 medRxiv
Top 0.1%
75.5%
Show abstract

Transmembrane protein 184B (TMEM184B) is an endosomal 7-pass transmembrane protein with evolutionarily conserved roles in synaptic structure and axon degeneration. We report six pediatric cases who have de novo heterozygous variants in TMEM184B; five individuals harbor a rare missense variant and one individual has an mRNA splice site change. This cohort is unified by overlapping neurodevelopmental deficits including developmental delay, corpus callosum hypoplasia, seizures, and/or microcephaly. TMEM184B is predicted to contain a pore domain wherein four of five human disease-associated missense variants cluster. Structural modeling suggests that all missense variants alter TMEM184B protein stability. To understand the contribution of TMEM184B to neural development in vivo, we knocked down the TMEM184B ortholog in zebrafish and observed microcephaly and reduced anterior commissural axons, aligning with symptoms of affected individuals. Ectopic expression of TMEM184B c.550A>G; p.Lys184Glu and c.484G>A; p.Gly162Arg variants cause reduced head size and body length, indicating dominant effects, while three other variants show haploinsufficiency. None of the variants are able to rescue the knockdown phenotype. Human induced pluripotent stem cells (iPSC) with monoallelic production of p.Lys184Glu show mRNA disruptions in key metabolic pathways including those controlling mechanistic target of rapamycin (mTOR) activity. Expression of p.Lys184Glu and c.863G>C; p.Gly288Ala increased apoptosis in cell lines and p.Lys184Glu increased nuclear localization of transcription factor EB (TFEB), consistent with a cellular starvation state. Together, our data indicate that TMEM184B variants cause cellular metabolic disruption and result in abnormal neural development.

2
Benchmarking autosomal recessive disease prevalence estimation from allele frequencies against newborn screening data

Sierant, M. C.; Knoblauch, N.; Witt, E.; Gaffney, D.; Pulit, S.; Wuster, A.

2025-10-13 genetic and genomic medicine 10.1101/2025.10.11.25337773 medRxiv
Top 0.1%
75.0%
Show abstract

Accurate estimates for the prevalence of rare congenital diseases are critical for understanding disease epidemiology and enabling drug development. Prevalence estimates can inform public health investment, identify communities with high disease burden or underdiagnosis, and reveal areas of unmet clinical need. With the advent of global-scale biobanks, genetics-based models to estimate the prevalence of disease have become viable. Autosomal recessive (AR) rare diseases are particularly tractable for this approach given that disease prevalence can be estimated from the pathogenic allele frequency (AF) in carriers from unaffected populations. Despite the usefulness of such estimates, this approach has not been validated against real-world clinical datasets at scale. Newborn screening (NBS) programs, which test newborns for a panel of neonatal diseases using quantitative diagnostic methods, provide a comparator for birth prevalence with low ascertainment bias, large sample size, and low diagnostic variability. NBS datasets thus offer a uniquely robust benchmark to evaluate and improve the accuracy of AR genetic prevalence models. Here we explore the feasibility, utility, and pitfalls of estimating AR birth prevalence using genetic and NBS data. We applied a genetic model to estimate birth prevalence for 28 AR diseases consistently present on NBS panels and benchmarked these against reported NBS birth prevalence in more than 12 million newborns in the United States. We found concordance between the genetic estimate and NBS was impacted by the population database used to derive AF, ancestry-matching methodology, and pathogenic variant inclusion criteria. Incorporating these refinements, we demonstrate that a genetics-first approach can provide order-of-magnitude estimates of AR disease birth prevalence for nearly all tested diseases (25/28; 89%). However, we note a general underestimate of the genetic prevalence, suggesting identifying additional pathogenic variants would improve the concordance with NBS. Further, we also assessed the impact of epidemiological and genetic variables, highlighting diseases where genetic prevalence estimates may not be suitable.

3
Long-read transcriptome analysis using IsoRanker for identifying pathogenic variants in Mendelian conditions

Cheng, Y.-H. H.; Sedeno-Cortes, A. E.; Ranchalis, J. E.; Munson, K. M.; Vollger, M. R.; Balton, E.; Genetti, C. A.; Undiagnosed Diseases Network, ; Genomics Research to Elucidate the Genetics of Rare Diseases consortium, ; University of Washington Center for Rare Diseases Research, ; Wojcik, M. H.; Beggs, A. H.; Bamshad, M. J.; Wei, C.-L.; Dipple, K. M.; Kumar, R. D.; Blue, E. E.; Jarvik, G.; Chong, J. X.; Witten, D. M.; O'Donnell-Luria, A.; Stergachis, A. B.

2025-11-13 genetic and genomic medicine 10.1101/2025.11.07.25339764 medRxiv
Top 0.1%
74.9%
Show abstract

Identifying pathogenic non-coding variants that contribute to Mendelian conditions remains challenging as the functional impact of these variants on gene function is often unknown. We present IsoRanker, a long-read transcriptome sequencing-based framework that prioritizes functionally relevant non-coding variants by detecting genes and novel isoforms with outlier expression, allelic imbalance, and/or nonsense-mediated decay (NMD). We generated paired cycloheximide-treated and untreated fibroblast transcriptomes from 31 individuals (3 individuals with known transcript-altering rare variants and 28 individuals with unsolved conditions) and linked transcripts to phased long-read genomes. IsoRanker successfully recovered known transcript alterations in this cohort and remained robust in subsampling analyses to cohorts of 11 individuals and [~]5 million full-length transcripts per individual. However, performance was dependent upon de novo isoform caller choice, particularly for NMD-sensitive and novel isoforms. Among 28 previously unsolved cases, IsoRanker deprioritized most fibroblast-expressed candidate splice site variants while nominating new leads. In one individual, IsoRanker prioritized HARS1, revealing biallelic non-coding variants that together produced a partial HARS1 loss-of-function and informed targeted therapy in this individual using histidine supplementation. These findings establish long-read, NMD-aware transcriptomics with IsoRanker as an effective approach for generating isoform-level functional evidence, improving classification of non-coding variants and supporting the diagnosis of individuals with rare diseases.

4
Individuals whose phenotype deviates from genetic expectation defined by common variation are enriched for rare damaging variants in genes that cause rare disease

Baya, N. A.; Lassen, F. H.; Hill, B.; Venkatesh, S. S.; Currant, H.; Lindgren, C. M.; Palmer, D. S.

2025-12-31 genetic and genomic medicine 10.64898/2025.12.30.25343229 medRxiv
Top 0.1%
59.0%
Show abstract

Polygenic scores (PGS) predict complex traits and stratify disease risk but often fail to fully capture individual-level variation. "Misaligned" individuals, whose observed phenotypes deviate from their genetically expected values based on polygenic scores (PGS), provide a powerful model for identifying factors beyond common-variant effects, including additional genetic factors. Here, we apply misalignment classification and enrichment testing frameworks to seven continuous and three dichotomous traits, assessing whether misaligned individuals in the UK Biobank are enriched for rare (minor allele frequency (MAF) < 0.1%) damaging genetic variation. We identify significant enrichment (false discovery rate (FDR)-adjusted P < 0.05) of predicted loss-of-function (pLoF) variants in COPB2 and GORAB among individuals misaligned for lower-than-expected bone mineral density. We refine previously observed grouped-gene enrichment in individuals with misaligned stature to the single-gene level: shorter-than-expected individuals are enriched for pLoF variants in ACAN and IGF1, and taller-than-expected individuals are enriched for predicted damaging missense in FBN1. Using an individuals misalignment classification as a phenotype, we perform an exome-wide scan across seven traits, resulting in 74 FDR- significant genes. We identify KANK1 as a gene associated with later age at menopause, potentially protective against primary ovarian insufficiency. For dichotomous disease status traits, we demonstrate evidence for the liability threshold model in the context of counteracting conditionally-orthogonal common and rare variant pathogenic/protective effects. Among individuals diagnosed with type 2 diabetes, carriers of rare pathogenic pLoF variants in HNF1A and HNF4A had significantly lower polygenic risk than non- carriers (FDR-adjusted one-sided t-test P < 5 x 10-3). We also show that coronary artery disease controls carrying rare protective pLoF variants in ANGPTL3 had nominally higher polygenic risk (one-sided t-test P = 0.03) than non-carriers. This study highlights the power of misalignment-based analyses in complex continuous phenotypes and disease, with the potential to validate known genetic contributors to traits and identify novel genes. This work paves the way for better molecular diagnoses and targeted therapeutic discovery.

5
SCAMPI: A scalable statistical framework for genome-wide interaction testing harnessing cross-trait correlations

Bian, S.; Bass, A. J.; Liu, Y.; Wingo, A. P.; Wingo, T.; Cutler, D. J.; Epstein, M. P.

2024-09-14 genetics 10.1101/2024.09.10.612314 medRxiv
Top 0.1%
58.1%
Show abstract

Family-based heritability estimates of complex traits are often considerably larger than their single-nucleotide polymorphism (SNP) heritability estimates. This discrepancy may be due to non-additive effects of genetic variation, including variation that interacts with other genes or environmental factors to influence the trait. Variance-based procedures provide a computationally efficient strategy to screen for SNPs with potential interaction effects without requiring the specification of the interacting variable. While valuable, such variance-based tests consider only a single trait and ignore likely pleiotropy among related traits that, if present, could improve power to detect such interaction effects. To fill this gap, we propose SCAMPI (Scalable Cauchy Aggregate test using Multiple Phenotypes to test Interactions), which screens for variants with interaction effects across multiple traits. SCAMPI is motivated by the observation that SNPs with pleiotropic interaction effects induce genotypic differences in the patterns of correlation among traits. By studying such patterns across genotype categories among multiple traits, we show that SCAMPI has improved performance over traditional univariate variance-based methods. Like those traditional variance-based tests, SCAMPI permits the screening of interaction effects without requiring the specification of the interaction variable and is further computationally scalable to biobank data. We employed SCAMPI to screen for interacting SNPs associated with four lipid-related traits in the UK Biobank and identified multiple gene regions missed by existing univariate variance-based tests. SCAMPI is implemented in software for public use.

6
Fine-Mapping and Credible Set Construction using a Multipopulation Joint Analysis of Marginal Summary Statistics from Genome-wide Association Studies

Shen, J.; Jiang, L.; Wang, K.; Wang, A.; Chen, F.; Newcombe, P. J.; Haiman, C. A.; Conti, D. V.

2022-12-22 genetics 10.1101/2022.12.22.521659 medRxiv
Top 0.1%
57.0%
Show abstract

Recent advancement in Genome-wide Association Studies (GWAS) comes from not only increasingly larger sample sizes but also the shifted focus towards underrepresented populations. Multi-population GWAS may increase power to detect novel risk variants and improve fine-mapping resolution by leveraging evidence from diverse populations and accounting for the difference in linkage disequilibrium (LD) across ethnic groups. Here, we expand upon our previous approach for single-population fine-mapping through Joint Analysis of Marginal SNP Effects (JAM) to a multi-population analysis (mJAM). Under the assumption that true causal variants are common across studies, we implement a novel version of JAM that conditions on multiple SNPs while explicitly incorporating the different LD structures across populations. The mJAM framework can be used to first select index variants using the mJAM likelihood with any feature selection approach. In addition, we present a novel approach leveraging the ideas of mediation to construct credible sets for these index variants. Construction of such credible sets can be performed given any existing index variants. We illustrate the implementation of the mJAM likelihood through two implementations: mJAM-SuSiE (a Bayesian approach) and mJAM-Forward selection. Through simulation studies based on realistic effect sizes and levels of LD, we demonstrated that mJAM performs better than other existing multi-ethnic methods for constructing concise credible sets that include the underlying causal variants. In real data examples taken from the most recent multi-population prostate cancer GWAS, we showed several practical advantages of mJAM over other existing methods.

7
Systematic comparison of phenome-wide admixture mapping and genome-wide association in a diverse biobank

Cullina, S.; Shemirani, R.; Asgari, S.; Kenny, E. E.

2024-11-18 genetic and genomic medicine 10.1101/2024.11.18.24317494 medRxiv
Top 0.1%
56.4%
Show abstract

Biobank-scale association studies that include Hispanic/Latino(a) (HL) and African American (AA) populations remain underrepresented, limiting the discovery of disease associated genetic factors in these groups. We present here a systematic comparison of phenome-wide admixture mapping (AM) and genome-wide association (GWAS) using data from the diverse BioMe biobank in New York City. Our analysis highlights 77 genome-wide significant AM signals, 48 of which were not detected by GWAS, emphasizing the complementary nature of these two approaches. AM-tagged variants show significantly higher minor allele frequency and population differentiation (Fst) while GWAS demonstrated higher odds ratios, underscoring the distinct genetic architecture identified by each method. This study offers a comprehensive phenome-wide AM resource, demonstrating its utility in uncovering novel genetic associations in underrepresented populations, particularly for variants missed by traditional GWAS approaches.

8
Estimation of Direct and Indirect Polygenic Effects and Gene-Environment Interactions using Polygenic Scores in Case-Parent Trio Studies

Wang, Z.; Grosvenor, L.; Ray, D.; Ruczinski, I.; Beaty, T.; Volk, H.; Ladd-Acosta, C.; Chatterjee, N.

2024-10-09 genetic and genomic medicine 10.1101/2024.10.08.24315066 medRxiv
Top 0.1%
55.5%
Show abstract

Family-based studies provide a unique opportunity to characterize genetic risks of diseases in the presence of population structure, assortative mating, and indirect genetic effects. We propose a novel framework, PGS-TRI, for the analysis of polygenic scores (PGS) in case-parent trio studies to estimate the risk of an index condition associated with direct effects of inherited PGS, indirect effects of parental PGS, and gene-environment interactions. Extensive simulation studies demonstrate the robustness of PGS-TRI in the presence of complex population structure and assortative mating. We applied PGS-TRI to multi-ancestry trio studies of autism spectrum disorders (ASD) (Ntrio = 18,383), deriving transmission-based estimates of risk for the direct effects of established PGS for ASD and other neurocognitive traits PGS across different ancestry groups and along a genetic ancestry continuum. Our analysis also identified significant indirect effects from parental PGS for BMI and several neurocognitive traits on childrens ASD risk. We also employed PGS-TRI in a trio study of European and Asian orofacial clefts (OFCs) (Ntrio = 1,904), investigating both direct and indirect effect of an established PGS and its interaction with maternal risk factors. Finally, we applied PGS-TRI to investigate the direct and indirect effects of large-scale transcriptome-wide and metabolome-wide traits on ASD and OFCs risks.

9
Common variant approaches to study Mendelian disease gene function identify novel phenome and pathways associated with PLOD3

Scalici, A.; Baker, J. T.; Blostein, F.; Shuey, M.; Choudhary, D.; Knapik, E. W.; Samuels, D. C.; Below, J. E.; Bastarache, L.; Miller-Fleming, T. W.; Cox, N. J.

2025-11-27 genetic and genomic medicine 10.1101/2025.11.25.25340832 medRxiv
Top 0.1%
54.4%
Show abstract

BackgroundThe study of rare and common genetic disorders, in terms of study design, methods, and their genetic architecture, has largely been thought of as distinct. As sequencing technologies and analysis methods have advanced, we have learned that polygenic background can affect the penetrance, severity, and onset of certain Mendelian conditions. While genome-wide analyses have significantly contributed to our understanding of human disease, these studies rarely explore how common variation within a traditionally Mendelian disease genes affects phenotypic outcomes. MethodsIn this study, we leverage common variant-based approaches and electronic health record (EHR)-derived phenome to study the phenotypic consequences associated with PLOD3, a Mendelian disease gene associated with BCARD syndrome. We conducted a gene-based phenome-wide association study (PheWAS) to identify phenotypes associated with reduced genetically predicted gene expression (GPGE) of PLOD3 in BioVU. To further quantify the phenotypic features associated with PLOD3, we leveraged a phenotype risk score (PheRS) constructed from the Mendelian features BCARD syndrome in OMIM and a PheRS constructed from our gene-based PheWAS. We used these PheRSs in a TWAS farmwork to expand genetic and pathway level associations. ResultsWe found that reduced GPGE of PLOD3 derived from common variants, can capture clinical phenome associated with Mendelian disease (BCARD syndrome) in addition to novel phenotypes not previously associated with PLOD3. These novel phenotypes were largely replicated in analyses of a protein quantitative trait loci (pQTL) for PLOD3 and in identity by descent (IBD) analysis of the PLOD3 locus. By generating PheRS for the Mendelian and gene-based phenome and using them in a TWAS framework, we identified novel genetic and pathway level associations. ConclusionsIn this study, we present a scalable approach using EHR-based phenome and GPGE to identify novel phenome and use it as a tool for gene discovery to identify novel gene and pathway associations. This study leveraged biobank scale genetic and phenotype data to identify candidate phenotypes for expansion of the PLOD3 phenome and identify potential novel disease mechanisms. The application of these approaches to other Mendelian disease genes has the potential to aid in drug repurposing and identify candidate therapeutic targets.

10
Deleterious, protein-altering variants in the X-linked transcriptional coregulator ZMYM3 in 22 individuals with a neurodevelopmental delay phenotype

Hiatt, S. M.; Trajkova, S.; Rossi Sebastiano, M.; Partridge, E. C.; Abidi, F. E.; Anderson, A.; Ansar, M.; Antonarakis, S. E.; Azadi, A.; Bachmann-Gagescu, R.; Bartuli, A.; Benech, C.; Berkowitz, J. L.; Betti, M. J.; Brusco, A.; Cannon, A.; Caron, G.; Chen, Y.; Crenshaw, M. M.; Cuisset, L.; Curry, C. J.; Darvish, H.; Demirdas, S.; Descartes, M.; Douglas, J.; Dyment, D. A.; Zghal Elloumi, H.; Ermondi, G.; Faoucher, M.; Farrow, E. G.; Felker, S. A.; Fisher, H.; Hurst, A. C. E.; Joset, P.; Kmoch, S.; Leadem, B. R.; Macchiaiolo, M.; Magner, M.; Mandrile, G.; Mattioli, F.; McEown, M.; Meadows, S. K

2022-09-30 genetic and genomic medicine 10.1101/2022.09.29.22279724 medRxiv
Top 0.1%
52.7%
Show abstract

Neurodevelopmental disorders (NDDs) often result from highly penetrant variation in one of many genes, including genes not yet characterized. Using the MatchMaker Exchange, we assembled a cohort of 22 individuals with rare, protein-altering variation in the X-linked transcriptional coregulator gene ZMYM3. Most (n=19) individuals were males; 15 males had maternally-inherited alleles, three of the variants in males arose de novo, and one had unknown inheritance. Overlapping features included developmental delay, intellectual disability, behavioral abnormalities, and a specific facial gestalt in a subset of males. Variants in almost all individuals (n=21) are missense, two of which are recurrent. Three unrelated males were identified with inherited variation at R441, a site at which variation has been previously reported in NDD-affected males, and two individuals have de novo variation at R1294. All variants affect evolutionarily conserved sites, and most are predicted to damage protein structure or function. ZMYM3 is relatively intolerant to variation in the general population, is highly expressed in the brain, and encodes a component of the KDM1A-RCOR1 chromatin-modifying complex. ChIP-seq experiments on one mutant, ZMYM3R1274W, indicate dramatically reduced genomic occupancy, supporting a hypomorphic effect. While we are unable to perform statistical evaluations to support a conclusive causative role for variation in ZMYM3 in disease, the totality of the evidence, including the presence of recurrent variation, overlapping phenotypic features, protein-modeling data, evolutionary constraint, and experimentally-confirmed functional effects, strongly supports ZMYM3 as a novel NDD gene.

11
A framework for integrated clinical risk assessment using population sequencing data

Fife, J. D.; Tran, T.; Bernatchez, J. R.; Shepard, K. E.; Koch, C.; Patel, A. P.; Fahed, A. C.; Krishnamurthy, S.; Genetics Center, R.; Collaboration, D.; Wang, W.; Buchanan, A. H.; Carey, D. J.; Metpally, R.; Khera, A. V.; Lebo, M.; Cassa, C. A.

2021-08-13 genetic and genomic medicine 10.1101/2021.08.12.21261563 medRxiv
Top 0.1%
52.6%
Show abstract

ImportanceClinical risk prediction for monogenic coding variants remains challenging even in established disease genes, as variants are often so rare that epidemiological assessment is not possible. These variants are collectively common in population cohorts -- one in six individuals carries a rare variant in nine clinically actionable genes commonly used in population health screening. ObjectiveTo expand diagnostic risk assessment in genomic medicine by integrating monogenic, polygenic, and clinical risk factors, and to classify individuals who carry monogenic variants as having elevated risk or population-level risk. Design, Setting, and ParticipantsParticipants aged 40-70 years were recruited from 22 UK assessment centers from 2006 to 2010. Monogenic, polygenic, and clinical risk factors are used to generate integrated predictions of risk for carriers of rare missense variants in 200,625 individuals with exome sequencing data. Relative risks and classification thresholds are validated using 92,455 participants in the Geisinger MyCode cohort recruited from 70 US sites from 2007 onward. Conclusions and RelevanceUsing integrated risk predictions, we identify 18.22% of UK Biobank (UKB) participants carrying variants of uncertain significance are at elevated risk for breast cancer (BC), familial hypercholesterolemia (FH), and colorectal cancer (CRC), accounting for 2.56% of the UKB in total. These predictions are concordant with clinical outcomes: individuals classified as having high risk have substantially higher risk ratios (Risk Ratio=3.71 [3.53, 3.90] BC, RR=4.71 [4.50, 4.92] FH, RR=2.65 [2.15, 3.14] CRC, logrank p<10-5), findings that are validated in an independent cohort ({chi}2 p=9.9x10-4 BC,{chi} 2 p=3.72x10-16 FH). Notably, we predict that 64% of UKB patients with laboratory-classified pathogenic FH variants are not at increased risk for coronary artery disease (CAD) when considering all patient and variant characteristics, and find no significant difference in CAD outcomes between these individuals and those without a monogenic disease-associated variant (logrank p=0.68). Current clinical practice guidelines discourage the disclosure of variants of uncertain significance to patients, but integrated modeling broadens this risk analysis, and identifies over 2.5-fold additional individuals who could potentially benefit from such information. This framework improves risk assessment within two similarly ascertained biobank cohorts, which may be useful in guiding preventative care and clinical management. Key PointsO_ST_ABSQuestionC_ST_ABSCan personalized risk assessments that consider monogenic, polygenic, and clinical characteristics improve diagnostic accuracy over traditional variant-level genetic assessments? FindingsIn established disease genes, we predict many carriers of variants of uncertain significance have significantly elevated risk. Conversely, we identify a substantial number of patients with known pathogenic coding variants who are unlikely to develop associated disorders. MeaningMany individuals would not learn about elevated risk for disease under current genetic diagnostic guidelines. Integrated risk assessments provide significant benefits over variant-only interpretation, and should be further evaluated for their potential to optimize clinical management, inform preventive care, and reduce potential harms.

12
Impact of cross-ancestry genetic architecture on GWAS in admixed populations

Mester, R.; Hou, K.; Ding, Y.; Meeks, G.; Burch, K. S.; Bhattacharya, A.; Henn, B. M.; Pasaniuc, B.

2023-01-24 bioinformatics 10.1101/2023.01.20.524946 medRxiv
Top 0.1%
52.3%
Show abstract

Genome-wide association studies (GWAS) have identified thousands of variants for disease risk. These studies have predominantly been conducted in individuals of European ancestries, which raises questions about their transferability to individuals of other ancestries. Of particular interest are admixed populations, usually defined as populations with recent ancestry from two or more continental sources. Admixed genomes contain segments of distinct ancestries that vary in composition across individuals in the population, allowing for the same allele to induce risk for disease on different ancestral backgrounds. This mosaicism raises unique challenges for GWAS in admixed populations, such as the need to correctly adjust for population stratification to balance type I error with statistical power. In this work we quantify the impact of differences in estimated allelic effect sizes for risk variants between ancestry backgrounds on association statistics. Specifically, while the possibility of estimated allelic effect-size heterogeneity by ancestry (HetLanc) can be modeled when performing GWAS in admixed populations, the extent of HetLanc needed to overcome the penalty from an additional degree of freedom in the association statistic has not been thoroughly quantified. Using extensive simulations of admixed genotypes and phenotypes we find that modeling HetLanc in its absence reduces statistical power by up to 72%. This finding is especially pronounced in the presence of allele frequency differentiation. We replicate simulation results using 4,327 African-European admixed genomes from the UK Biobank for 12 traits to find that for most significant SNPs HetLanc is not large enough for GWAS to benefit from modeling heterogeneity.

13
LANTERN: Leveraging Local Ancestry Tracts to Enhance Rare-Variant Aggregate Association Testing

Wang, Y.; Tuftin, B.; Raffield, L. M.; Hidalgo, B.; Kerns, S. L.; DeWan, A. T.; Leal, S. M.; Auer, P.

2026-04-27 genetic and genomic medicine 10.64898/2026.04.24.26351693 medRxiv
Top 0.1%
51.8%
Show abstract

Individuals with admixed ancestry comprise a significant proportion of populations of the Americas. Statistical methods have been developed to specifically leverage local ancestry inference to enhance the power and interpretability of genome-wide association studies in admixed populations. However, no such methods currently exist to test for rare-variant aggregate associations. Here we present LANTERN (Leveraging local ANcestry Tracts to Enhance Rare variaNt aggregate associations), a method that infers the alleles that lie on each ancestral haplotype and conducts rare-variant aggregate association testing in a generalized linear mixed model framework. Through simulation studies we demonstrated that LANTERN achieves proper control of Type 1 error while boosting power to detect associations when causal alleles predominately lie on one ancestral haplotype. Using data from a cohort of African American participants from the Jackson Heart Study, LANTERN identified two genes known to be involved in red-blood cell (RBC) biology when local ancestry information was incorporated. Specifically, a burden of rare alleles on European ancestral haplotypes in EPO was associated with both hemoglobin levels (HGB) and RBC counts, whereas a burden of rare alleles on African ancestral haplotypes in EPB42 was associated with HGB and RBC. In summary, LANTERN (i) allows for the identification of ancestry-specific rare-variant associations; and (ii) enhances rare-variant association signals compared to an analysis that ignores local ancestry. LANTERN is implemented in R and is freely available on GitHub.

14
Mind the gap: characterizing bias due to population mismatch in two-sample Mendelian randomization

Li, J.; Morrison, J.

2025-08-01 epidemiology 10.1101/2025.07.30.25332465 medRxiv
Top 0.1%
51.6%
Show abstract

1Mendelian randomization (MR) is a statistical method for estimating causal effects using genetic variants as instrumental variables. In two sample MR (2SMR), different study samples are used to estimate genetic associations with the exposure and outcome. For valid inference, these studies must include individuals from the same population. Using studies from different populations may bias the MR estimate due to differences in variant-exposure associations resulting from differences in linkage disequilibrium or genetic effects on the exposure trait. We show that violation of the same-population assumption leads to bias in the causal estimate towards zero on average, and does not increase the rate of false positives when using the most common MR study design. We verify this result in a broad survey of MR estimates, comparing estimates made with matching and mismatching populations across 546 trait pairs measured in 2-7 ancestries. We find that most population-mismatched estimates are attenuated towards zero compared to their corresponding population-matched estimates, and that increasing genetic distance between study populations is associated with greater shrinkage. We observe bias even when mismatched populations have the same continental ancestry. However, we also find that, in some cases, using a larger exposure study with mismatching ancestry can improve power by dramatically increasing precision. These results show that even intra-continental population mismatch can bias MR estimates, but also suggests there is potential to improve the power of MR in understudied populations by properly leveraging larger, mismatching study populations.

15
Integrating Whole Genome and Transcriptome Sequencing to Characterize the Genetic Architecture of Isoform Variation and its Implications for Health and Disease

Liu, C.; Joehanes, R.; Ma, J.; Xie, J.; Yang, J.; Wang, M.; Huan, T.; Hwang, S.-J.; Wen, J.; Sun, Q.; Demirkale, C.; Heard-Costa, N.; Orchar, P.; Carson, A. P.; Raffield, L. M.; Reiner, A. P.; Li, Y.; O'Connor, G.; Murabito, J. M.; Munson, P.; Levy, D.

2024-12-06 epidemiology 10.1101/2024.12.04.24318434 medRxiv
Top 0.1%
49.7%
Show abstract

We created a comprehensive whole blood splice variation quantitative trait locus (sQTL) resource by analyzing isoform expression ratio (isoform-to-gene) in Framingham Heart Study (FHS) participants (discovery: n=2,622; validation: n=1,094) with whole genome (WGS) and transcriptome sequencing (RNA-seq) data. External replication was conducted using WGS and RNA-seq from the Jackson Heart Study (JHS, n=1,020). We identified over 3.5 million cis-sQTL-isoform pairs (p<5e-8), comprising 1,176,624 cis-sQTL variants and 10,883 isoform transcripts from 4,971 sGenes, with significant change in isoform-to-gene ratio due to allelic variation. We validated 61% of these pairs in the FHS validation sample (p<1e-4). External validation (p<1e-4) in JHS for the top 10,000 and 100,000 most significant cis-sQTL-isoform pairs was 88% and 69%, respectively, while overall pairs validated at 23%. For 20% of cis-sQTLs in the FHS discovery sample, allelic variation did not significantly correlate with overall gene expression. sQTLs are enriched in splice donor and acceptor sites, as well as in GWAS SNPs, methylation QTLs, and protein QTLs. We detailed several sentinel cis-sQTLs influencing alternative splicing, with potential causal effects on cardiovascular disease risk. Notably, rs12898397 (T>C) affects splicing of ULK3, lowering levels of the full-length transcript ENST00000440863.7 and increasing levels of the truncated transcript ENST00000569437.5, encoding proteins of different lengths. Mendelian randomization analysis demonstrated that a lower ratio of the full-length isoform is causally associated with lower diastolic blood pressure and reduced lymphocyte percentages. This sQTL resource provides valuable insights into how transcriptomic variation may influence health outcomes.

16
Structural variant discovery and diagnostic impact in rare diseases from short-read and long-read sequencing

Sanchis-Juan, A.; Mostovoy, Y.; Stenton, S. L.; Ganesh, V. S.; Weisburd, B.; Yenkin, A.; Kurtas, N. E.; Zhao, X.; Shin, E.; Boone, P. M.; Su, H.; Lee, A. S.; Yadav, R.; Allan, K.; Argilli, E.; Austin-Tse, C.; Barry, B. J.; Baxter, S.; Beggs, A. H.; Bell, K. M.; Blankenmeister, B.; Bönnemann, C. G.; Brownstein, C. A.; Bujakowska, K. M.; Carbonell, E.; Cooper, S. T.; Covill, L. E.; DiTroia, S.; Donkervoort, S.; Engle, E. C.; Gallacher, L.; Genetti, C. A.; Gleeson, J. G.; Guan, B.; Hall, S.; Hildebrandt, F.; Hufnagel, R. B.; Jurgens, J. A.; Khorgade, A.; Lemire, G.; Liau, E.; Ma, J.; Madden, J.

2026-06-24 genetic and genomic medicine 10.64898/2026.06.22.26356238 medRxiv
Top 0.1%
48.9%
Show abstract

Rare diseases collectively affect 1 in 10 individuals, yet current genetic testing fails to identify a causal variant for most cases. At present, cytogenetic methods and/or sequencing approaches such as exome (ES) or short-read genome sequencing (srGS) represent the state-of-the-art for comprehensive clinical discovery of sequence and structural variants (SVs), including copy number variants, balanced SVs, complex SVs, and tandem repeats (TRs). Recently, long-read genome sequencing (lrGS), coupled with multiomics data, has presented great promise to resolve variation in genomic regions recalcitrant to characterization by srGS such as highly repetitive simple repeat sequences and segmental duplications. However, there are few guidelines to enable clinical interpretation of genetic variation in these highly repetitive genomic regions, and the enthusiasm of the field in adopting lrGS has made it difficult to assess the true added diagnostic yield of this technology due to widely variable and inconsistently applied analytic pipelines and variable degrees of pre-screening by ES or srGS. Here, we investigated the contribution of SVs to rare diseases using srGS as a front-line strategy when paired with highly sensitive SV discovery and evaluate the added diagnostic yield of incorporating lrGS for a subset of cases. Our srGS analysis encompassed 1,462 families (3,450 individuals) recruited through the Broad Institute Center for Mendelian Genetics and the Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) programs. Diagnostic SVs were identified in 5.4% of cases (79/1,462), of which 80% were uniquely detectable by srGS compared to standard cytogenetic techniques. For 96 families (including 10 families with a heterozygous variant observed in a known recessive gene of clinical relevance), we performed lrGS with methylation profiling, as well as long-read transcriptomic analyses in a subset of 20 trios. Analyses with lrGS yielded over 25,000 SVs per genome, 63% of which were not captured by srGS, along with an additional ~200 rare SNV/indels per genome not previously captured and 12 differentially methylated regions per genome. Among these, we identified only one diagnostic variant not interpreted by srGS, an apparently mosaic de novo SNV in CASK that was absent in the srGS callset due to allelic imbalance. No new diagnoses were supported by long-read transcriptomics or episignatures. In this well characterized rare disease cohort, the added diagnostic yield was thus 1.04% (1/96 families). Following a systematic literature review of prior lrGS studies, we find that most reported diagnoses were detectable by srGS and that our added diagnostic yield is consistent with those prior studies. These studies emphasize the significant impact of comprehensive SV discovery in rare disease cases and further demonstrate the power for increased discovery of novel genomic variation and episignatures from lrGS. Nonetheless, they also serve to temper expectations of dramatic diagnostic advances in rare disease patients until there is more extensive annotation of the functional and clinical impact of all coding and noncoding variation uniquely accessible to lrGS with extensive reference databases spanning highly repetitive genomic sequencing that could be enabled by this transformative technology.

17
Systematic comparison of colocalization methods using protein quantitative trait loci

Hwang, S.; Pullin, J. M.; Wallace, C.; Whittaker, J.; Burgess, S.

2025-11-07 genetics 10.1101/2025.11.07.686776 medRxiv
Top 0.1%
48.5%
Show abstract

Colocalization is frequently performed as a step to triage findings from genetic investigations linking molecular and disease data. However, the reliability and consistency of the various colocalization methods is not well understood. In particular, it is unclear whether a non-colocalization result should be taken as a definitive sign that two traits do not share a genetic cause in the region of interest, or merely as suggestive evidence. We use protein QTL data to benchmark four colocalization methods (coloc, coloc-SuSiE, prop-coloc, and colocPropTest), considering whether the methods conclude there is colocalization between datasets representing associations with the same protein. We consider a baseline scenario in which the associations come from the same dataset split in half at random, and scenarios in which the associations are with the same protein, but measured on different platforms or in different populations. In the baseline scenario, all methods report colocalization for the majority of proteins. In other scenarios, methods do not consistently report colocalization, they often report non-colocalization, and they often disagree. In the worst-case scenario, colocalization was only agreed by all four methods for 20% of proteins, despite our experiment being constructed to select for cases where colocalization is expected. This suggests caution in performing and interpreting colocalization analyses is warranted.

18
Estimating Disorder Probability Based on Polygenic Prediction Using the BPC Approach

Uffelmann, E.; Major Depressive Disorder Working Group of the Psychiatric Genomics Consortium, ; Schizophrenia Working Group of the Psychiatric Genomics Consortium, ; Price, A. L.; Posthuma, D.; Peyrot, W. J.

2024-01-13 genetic and genomic medicine 10.1101/2024.01.12.24301157 medRxiv
Top 0.1%
48.4%
Show abstract

Polygenic Scores (PGSs) summarize an individuals genetic propensity for a given trait in a single value based on SNP effect sizes derived from Genome-Wide Association Study (GWAS) results. Methods have been developed that apply Bayesian approaches to improve the prediction accuracy of PGSs through optimization of estimated effect sizes. While these methods are generally well-calibrated for continuous traits (implying the predicted values are, on average, equal to the true trait values), they are not well-calibrated for binary disorder traits in ascertained samles. This is a problem because well-calibrated PGSs are needed to reliably compute the absolute disorder probability for an individual to facilitate future clinical implementation. Here, we introduce the Bayesian polygenic score Probability Conversion (BPC) approach, which computes an individuals predicted disorder probability using GWAS summary statistics, an existing Bayesian PGS method (e.g., PRScs, SBayesR), the individuals genotype data, and a prior disorder probability (which can be specified flexibly, based on e.g., literature, small reference samples, or prior elicitation). The BPC approach transforms the PGS to its underlying liability scale, computes the variances of the PGS in cases and controls, and applies Bayes Theorem to compute the absolute disorder probability; it is practical in its application as it does not require a tuning sample with both genotype and phenotype data. We applied the BPC approach to extensive simulated data and empirical data of nine disorders. The BPC approach yielded well-calibrated results that were consistently better than the results of another recently published approach.

19
De novo variants in KDM2A cause a syndromic neurodevelopmental disorder

Anderson, E. N.; Drukewitz, S. H.; Kour, S.; Chimata, A. V.; Rajan, D. S.; Schoennagel, S.; Stals, K. L.; Donnelly, D.; O'Sullivan, S.; Mantovani, J. F.; Tan, T. Y.; Stark, Z.; Zacher, P.; Chatron, N.; Monin, P.; Drunat, S.; Vial, Y.; Latypova, X.; Levy, J.; Verloes, A.; Carter, J. N.; Bonner, D. E.; Shankar, S. P.; Bernstein, J. A.; Cohen, J. S.; Comi, A.; Alexis Carere, D.; Dyer, L. M.; Mullegama, S. V.; Sanchez-Lara, P. A.; Grand, K.; Kim, H.-G.; Ben-Mahmoud, A.; Gospe, S. M.; Belles, R. S.; Bellus, G.; Lichtenbelt, K. D.; Oegema, R.; Rauch, A.; Ivanovski, I.; Tran-Mau-Them, F.; Garde, A.;

2025-04-02 genetic and genomic medicine 10.1101/2025.03.31.25324695 medRxiv
Top 0.1%
46.8%
Show abstract

Germline variants that disrupt components of the epigenetic machinery cause syndromic neurodevelopmental disorders. Using exome and genome sequencing, we identified de novo variants in KDM2A, a lysine demethylase crucial for embryonic development, in 18 individuals with developmental delays and/or intellectual disabilities. The severity ranged from learning disabilities to severe intellectual disability. Other core symptoms included feeding difficulties, growth issues such as intrauterine growth restriction, short stature and microcephaly as well as recurrent facial features like epicanthic folds, upslanted palpebral fissures, thin lips, and low-set ears. Expression of human disease-causing KDM2A variants in a Drosophila melanogaster model led to neural degeneration, motor defects, and reduced lifespan. Interestingly, pathogenic variants in KDM2A affected physiological attributes including subcellular distribution, expression and stability in human cells. Genetic epistasis experiments indicated that KDM2A variants likely exert their effects through a potential gain-of-function mechanism, as eliminating endogenous KDM2A in Drosophila did not produce noticeable neurodevelopmental phenotypes. Data from Enzymatic-Methylation sequencing supports the suggested gene-disease association by showing an aberrant methylome profiles in affected individuals peripheral blood. Combining our genetic, phenotypic and functional findings, we establish de novo variants in KDM2A as causative for a syndromic neurodevelopmental disorder.

20
Heterozygous loss-of-function SMC3 variants are associated with variable and incompletely penetrant growth and developmental features

Ansari, M.; Faour, K. N. W.; Shimamura, A.; Grimes, G.; Kao, E. M.; Denhoff, E. R.; Blatnik, A.; Ben-Isvy, D.; Wang, L.; Helm, B. M.; Firth, H.; Breman, A. M.; Bijlsma, E. K.; Iwata-Otsubo, A.; de Ravel, T. J. L.; Fusaro, V.; Fryer, A.; Nykamp, K.; Stuhn, L. G.; Haack, T. B.; Korenke, G. C.; Constantinou, P.; Bujakowska, K. M.; Low, K. J.; Place, E.; Humberson, J.; Napier, M. P.; Hoffman, J.; Juusola, J.; Deardorff, M. A.; Shao, W.; Rockowitz, S.; Krantz, I.; Kaur, M.; Raible, S.; Kliesch, S.; Singer-Berk, M.; Groopman, E.; DiTroia, S.; Ballal, S.; Srivastava, S.; Rothfelder, K.; Biskup, S.; R

2023-09-28 genetic and genomic medicine 10.1101/2023.09.27.23294269 medRxiv
Top 0.1%
46.5%
Show abstract

Heterozygous missense variants and in-frame indels in SMC3 are a cause of Cornelia de Lange syndrome (CdLS), marked by intellectual disability, growth deficiency, and dysmorphism, via an apparent dominant-negative mechanism. However, the spectrum of manifestations associated with SMC3 loss-of-function variants has not been reported, leading to hypotheses of alternative phenotypes or even developmental lethality. We used matchmaking servers, patient registries, and other resources to identify individuals with heterozygous, predicted loss-of-function (pLoF) variants in SMC3, and analyzed population databases to characterize mutational intolerance in this gene. Here, we show that SMC3 behaves as an archetypal haploinsufficient gene: it is highly constrained against pLoF variants, strongly depleted for missense variants, and pLoF variants are associated with a range of developmental phenotypes. Among 13 individuals with SMC3 pLoF variants, phenotypes were variable but coalesced on low growth parameters, developmental delay/intellectual disability, and dysmorphism reminiscent of atypical CdLS. Comparisons to individuals with SMC3 missense/in-frame indel variants demonstrated a milder presentation in pLoF carriers. Furthermore, several individuals harboring pLoF variants in SMC3 were nonpenetrant for growth, developmental, and/or dysmorphic features, some instead having intriguing symptomatologies with rational biological links to SMC3 including bone marrow failure, acute myeloid leukemia, and Coats retinal vasculopathy. Analyses of transcriptomic and epigenetic data suggest that SMC3 pLoF variants reduce SMC3 expression but do not result in a blood DNA methylation signature clustering with that of CdLS, and that the global transcriptional signature of SMC3 loss is model-dependent. Our finding of substantial population-scale LoF intolerance in concert with variable penetrance in subjects with SMC3 pLoF variants expands the scope of cohesinopathies, informs on their allelic architecture, and suggests the existence of additional clearly LoF-constrained genes whose disease links will be confirmed only by multi-layered genomic data paired with careful phenotyping.