Back

Personalized Feature Statistics: Individual-Level Variant Inference under Genetic Ancestry Continuum

Wang, J. F.; Yu, R.; Edelson, J.; Park, J.; Le Guen, Y.; Liu, X.; Belloy, M.; Ionita-Laza, I.; Greicius, M.; Tang, H.; He, Z.

2026-04-29 neurology
10.64898/2026.04.28.26351879 medRxiv
Show abstract

Genome-wide association studies (GWAS) have successfully identified numerous genetic variants associated with complex diseases. However, the extent to which the effects of these variants vary across populations of diverse ancestries remains poorly understood. Furthermore, in these contexts genetic ancestry is treated as a categorical variable, thereby oversimplifying its continuous nature and the more nuanced ways in which it can influence genetic effects on disease. Here, we propose personalized feature statistics (PFstatistics), a statistical framework that quantifies the importance of genetic variants to a phenotype based on each individuals ancestry background, and profiles heterogeneous genetic effects across the genetic ancestry continuum. We demonstrate the utility of this framework through both simulations and real data analysis using sequencing data from ancestrally diverse cohorts in the Alzheimers Disease Sequencing Project (ADSP). We show that Alzheimers Disease (AD) risk variants span a spectrum from ancestry-homogeneous to ancestry-dependent effects, and that PFstatistics characterizes this spectrum at individual resolution across the ancestry continuum. PFstatistics also provides individual-level variant selection with FDR controlled at a target level, yielding distinct selection sets that vary across individuals according to their ancestry background. While demonstrated in the context of genetic ancestry, the proposed method is broadly applicable to other heterogeneity features such as environmental factors, offering a robust tool for understanding complex genetic contributions across diverse populations.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Nature Communications
5641 papers in training set
Top 7%
19.2%
2
Human Genetics and Genomics Advances
84 papers in training set
Top 0.1%
12.3%
3
Nature Genetics
286 papers in training set
Top 0.5%
11.0%
4
Nature Computational Science
55 papers in training set
Top 0.1%
10.1%
50% of probability mass above
5
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 7%
6.5%
6
Nature Medicine
125 papers in training set
Top 0.5%
4.2%
7
Communications Biology
993 papers in training set
Top 8%
2.5%
8
Alzheimer's & Dementia
163 papers in training set
Top 1%
2.0%
9
PLOS Computational Biology
1863 papers in training set
Top 14%
1.8%
10
Scientific Reports
3612 papers in training set
Top 52%
1.8%
11
Advanced Science
286 papers in training set
Top 4%
1.7%
12
Genome Medicine
183 papers in training set
Top 3%
1.6%
13
Briefings in Bioinformatics
354 papers in training set
Top 5%
1.4%
14
Communications Medicine
113 papers in training set
Top 3%
1.2%
15
Neuron
337 papers in training set
Top 4%
1.2%
16
Nature Neuroscience
252 papers in training set
Top 4%
1.2%
17
The Journal of Immunology
166 papers in training set
Top 2%
1.2%
18
PLOS ONE
5266 papers in training set
Top 58%
1.0%
19
The American Journal of Human Genetics
234 papers in training set
Top 3%
0.9%
20
Nucleic Acids Research
1281 papers in training set
Top 13%
0.9%
21
Nature Biomedical Engineering
47 papers in training set
Top 1%
0.9%
22
Genome Biology
637 papers in training set
Top 9%
0.6%
23
Cell Systems
201 papers in training set
Top 6%
0.5%
24
Bioinformatics
1204 papers in training set
Top 10%
0.5%
25
npj Systems Biology and Applications
125 papers in training set
Top 3%
0.5%
26
NeuroImage
903 papers in training set
Top 7%
0.5%
27
eLife
5828 papers in training set
Top 71%
0.5%
28
Genetic Epidemiology
55 papers in training set
Top 0.9%
0.5%