Non-Parametric Ancestry Adjustment for Polygenic Scores
Mas Montserrat, D.; Barrabes, M.; Bustamante, C. D.; Ioannidis, A. G.
Show abstract
Modern polygenic risk scores (PRS) exhibit shifts correlated with ancestry, leading to erroneous predictions for non-European individuals when models are trained on predominantly European cohorts. Such shifts arise from, among other factors, (1) algorithmic limitations in the ability of PRS model training to detect causal variants, rather than nearby variants with ancestry-dependent correlations to the causal one, (2) under-representation of alleles with higher prevalence in non-European populations in the association study training, and (3) gene-by-environment interactions where the environment is correlated with genetic ancestry. Current ancestry-adjustment methodologies often discretize individuals into population categories and apply a simple affine mapping to reduce these genetic ancestry biases. However, such approaches provide suboptimal adjustments, particularly for admixed individuals. In this work, we introduce a detailed theoretical characterization of ancestry-dependent biases and propose novel methods based on non-parametric neighborhood techniques that provide more accurate empirical results and admit statistical consistency guarantees. Extensive experiments using the UK Biobank demonstrate the effectiveness of the proposed methods.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Scaling the Discrete-time Wright Fisher model to biobank-scale datasets 96%
- Reflection Knockoffs via Householder Reflection: Applications in Proteomics and Genetic Fine Mapping 96%
- Private Genomes and Public SNPs: Homomorphic encryption of genotypes and phenotypes for shared quantitative genetics 96%
Similar papers in this journal
- Fast Lasso method for Large-scale and Ultrahigh-dimensional Cox Model with applications to UK Biobank 95%
- Survival Analysis on Rare Events Using Group-Regularized Multi-Response Cox Regression 95%
- Estimating the overall fraction of phenotypic variance attributed to high-dimensional predictors measured with error 95%
Similar papers in this journal
- Probabilistic Cause-of-disease Assignment using Case-control Diagnostic Tests: A Latent Variable Regression Approach 95%
- Penalized reduced rank regression for multi-outcome survival data supports a common metabolic risk score for age-related diseases 95%
- Bayesian Variable Selection with a Pleiotropic Loss Function in Mendelian Randomization 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.