Unified meta regression models for rare variant association studies
Lauer, L.; Rivas, M. A.
Show abstract
Rare variant association studies (RVAS) of complex traits have emerged as a powerful approach to advance drug discovery and diagnostics. Missense pathogenicity predictions from AlphaMissense based on structural context and protein language models improve the differentiation between benign and deleterious variants. Constraint metrics, on the other hand, allow researchers to pinpoint genomic regions under selective pressure that may not directly impact protein structure, but are more likely to contain functionally important mutations. Loss-of-function (LoF) variants, which result in the complete or partial loss of protein function, are particularly informative, as it is more straightforward to assess their downstream functional consequences. In this study, we present a unified meta regression model approach that incorporates the probability of pathogenicity, probability of constraint, and indicator whether a variant is a predicted loss-of-function or missense variant as features to model the observed effect size and uncertainty of effect size obtained from single-variant genetic analysis. We applied the unified meta regression model to 1,144 continuous phenotypes from UK Biobank using single variant summary statistics obtained from Genebass. We replicated our findings using the AllofUS cohort. For each gene discovery, we make available a characterization of whether constrained sites are associated with the phenotype, whether pathogenic sites determined by structural based predictions are associated with phenotype, and whether broader loss-of-function or missense variant annotation better explains the summary statistics observed. Our results are publicly available at Global Biobank Engine (https://biobankengine.shinyapps.io/phenome-wide-unified-model/).
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Incorporating family disease history and controlling case-control imbalance for population based genetic association studies 95%
- acmgscaler: An R package and Colab for standardised gene-level variant effect score calibration within the ACMG/AMP framework 95%
- An exact, unifying framework for region-based association testing in family-based designs, including higher criticism approaches, SKATs, multivariate and burden tests 94%
Similar papers in this journal
- Scalable generalized linear mixed model for region-based association tests in large biobanks and cohorts 96%
- LDAK-KVIK performs fast and powerful mixed-model association analysis of quantitative and binary phenotypes 95%
- Computationally efficient whole genome regression for quantitative and binary traits 95%
Similar papers in this journal
- Ancestry adjustment improves genome-wide estimates of regional intolerance 95%
- Hierarchical clustering of gene-level association statistics reveals shared and differential genetic architecture among traits in the UK Biobank 93%
- Characterization of direct and/or indirect genetic associations for multiple traits in longitudinal studies of disease progression 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.