Back

Evaluating sequence-to-function deep learning models for ancestry-stratified regulatory variant effect prediction using multi-ancestry blood eQTLs

Sun, X.; Mews, M.; Wheeler, N. R.; Benchek, P.; Gu, T.; Gomez, L.; Mustafa, Y.; Wang, L.-S.; Leung, Y. Y.; Schellenberg, G. D.; Pericak-Vance, M. A.; Haines, J. L.; Griswold, A. J.; Bush, W. S.

2026-06-26 bioinformatics
10.64898/2026.06.22.730889 bioRxiv
Show abstract

Background: Sequence-to-function (S2F) deep learning models are increasingly used to prioritize non-coding regulatory variants, but their behavior across ancestrally diverse populations remains unclear. Because both training data and reference resources are heavily European-centered, multi-ancestry benchmarks are needed to determine whether S2F scores capture regulatory effects consistently across populations with different allele-frequency and LD patterns. Methods: We evaluated Borzoi and AlphaGenome using whole blood eQTL data from the MAGENTA cohort, including African American (AA; N=224), Caribbean Hispanic (CH; N=209), and Non-Hispanic White (NHW; N=235) participants. Model predictions were benchmarked against sampled nominal eQTLs and ancestry-stratified SuSiE fine-mapped variants using Spearman correlation, direction concordance, inter-model convergence, and distance-matched AUROC, with sensitivity analyses for minor allele frequency and comparison-set definition. We also compared FILER functional annotation overlap among high-Posterior Inclusion Probability (PIP) variants across ancestries. Results: Both models showed weak agreement with nominal eQTL effect sizes across ancestries and TSS-distance bins ({rho}[&le;]0.138), with direction concordance only marginally above chance. Agreement and discrimination improved for high-confidence fine-mapped variants, and Borzoi and AlphaGenome showed stronger inter-model convergence on fine-mapped variants than on nominal eQTLs, consistent with enrichment for regulatory variants whose effects are more apparent to sequence-based models. In distance-matched AUROC analyses at PIP [&ge;]0.9 using PIP <0.01 variants as low-PIP comparison variants, the AA high-PIP variant set yielded the highest discrimination for both Borzoi (0.837 [95% CI: 0.790-0.870]) and AlphaGenome (0.820 [0.793-0.845]). The CH-versus-NHW ordering was model-dependent: Borzoi yielded higher AUROC in NHW than CH, whereas AlphaGenome produced nearly identical CH and NHW estimates. AUROC values were lower when intermediate-PIP variants were used as comparison variants, but the AA set retained the highest discrimination. MAF-stratified sensitivity analyses attenuated some ancestry contrasts but did not eliminate the higher AA discrimination pattern. Functional annotation analysis showed that AA high-PIP variants more often overlapped chromatin accessibility and chromatin-contact annotations than NHW variants, despite lower overlap with prior eQTL and sQTL annotation catalogs. Conclusions: Borzoi and AlphaGenome showed limited agreement with nominal eQTL effect sizes, but better distinguished high-confidence fine-mapped eQTLs from low-PIP variants. These results support using S2F scores as prioritization evidence for fine-mapped regulatory variants, especially promoter-proximal high-PIP variants, rather than as standalone predictors of eQTL effect size. The strongest discrimination was observed for the AA high-PIP variant set. Overall, the AA result is best interpreted as stronger separation of high-PIP variants from lower-PIP comparison variants, shaped by fine-mapping resolution, LD, the choice of comparison variants, and annotation composition.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Human Genetics and Genomics Advances
84 papers in training set
Top 0.1%
28.3%
2
The American Journal of Human Genetics
234 papers in training set
Top 0.4%
10.9%
3
Genome Biology
637 papers in training set
Top 1%
7.8%
4
BMC Genomics
406 papers in training set
Top 1%
5.4%
50% of probability mass above
5
Nature Communications
5641 papers in training set
Top 30%
4.8%
6
BMC Bioinformatics
457 papers in training set
Top 2%
4.8%
7
Bioinformatics
1204 papers in training set
Top 5%
4.0%
8
Genes
144 papers in training set
Top 1%
2.4%
9
Briefings in Bioinformatics
354 papers in training set
Top 3%
2.4%
10
Frontiers in Genetics
230 papers in training set
Top 2%
2.1%
11
NAR Genomics and Bioinformatics
242 papers in training set
Top 2%
1.9%
12
Genome Medicine
183 papers in training set
Top 3%
1.7%
13
GENETICS
483 papers in training set
Top 3%
1.7%
14
Scientific Reports
3612 papers in training set
Top 59%
1.5%
15
Nature Genetics
286 papers in training set
Top 3%
1.5%
16
Computational and Structural Biotechnology Journal
242 papers in training set
Top 4%
1.3%
17
eLife
5828 papers in training set
Top 58%
1.1%
18
Bioinformatics Advances
203 papers in training set
Top 4%
1.1%
19
Nucleic Acids Research
1281 papers in training set
Top 11%
1.1%
20
Genome Research
468 papers in training set
Top 6%
1.0%
21
European Journal of Human Genetics
58 papers in training set
Top 1%
0.8%
22
npj Genomic Medicine
36 papers in training set
Top 0.8%
0.8%
23
PLOS Computational Biology
1863 papers in training set
Top 22%
0.6%
24
PLOS ONE
5266 papers in training set
Top 65%
0.6%
25
Human Molecular Genetics
141 papers in training set
Top 4%
0.6%