Back

Empirical estimation of multiple-testing burden for population-based HLA association studies using sequencing-derived HLA alleles across genetic ancestries

Taliun, D.; Gagliano Taliun, S. A.

2026-07-15 genetics
10.64898/2026.07.12.738059 bioRxiv
Show abstract

As population-scale whole-genome sequencing datasets continue to expand, they enable genetic association studies beyond single-nucleotide variants to more complex forms of genetic variation, including classical human leukocyte antigen (HLA) alleles. The HLA region comprises nine highly polymorphic classical HLA genes in extensive linkage disequilibrium that are associated with numerous autoimmune and infectious diseases. However, unlike genome-wide association studies of single-nucleotide variants, there is no general guidance for controlling the multiple-testing burden in HLA allele association analyses. Here, we systematically evaluated the effective number of independent HLA allele tests using sequencing data from diverse genetic ancestries, analytical derivation and simulations. We show that the multiple-testing burden depends on genetic ancestry, allele frequency, and the phenotype model, but remains remarkably stable across minor allele count thresholds, corresponding to approximately 60-70% of the total number of tested HLA alleles. Simulations further demonstrate that the effective number of tests can exceed 90% under realistic disease models. Analyses of 4-field HLA alleles from long-read sequencing showed that higher typing resolution increases the number of alleles but preserves the underlying correlation structure and scales the effective number of independent tests proportionally. Our results provide practical guidance for HLA association studies and support Bonferroni correction based on the total number of tested HLA alleles as a simple and robust approximation when permutation-based approaches are impractical.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
The American Journal of Human Genetics
234 papers in training set
Top 0.1%
30.7%
2
Nature Communications
5641 papers in training set
Top 18%
9.7%
3
Human Genetics and Genomics Advances
84 papers in training set
Top 0.1%
8.8%
4
Nature Genetics
286 papers in training set
Top 0.8%
7.8%
50% of probability mass above
5
GENETICS
483 papers in training set
Top 0.7%
7.8%
6
PLOS Genetics
862 papers in training set
Top 3%
4.8%
7
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 12%
4.3%
8
G3: Genes|Genomes|Genetics
35 papers in training set
Top 0.1%
2.4%
9
eLife
5828 papers in training set
Top 45%
2.1%
10
Nature Computational Science
55 papers in training set
Top 0.8%
1.4%
11
Frontiers in Genetics
230 papers in training set
Top 3%
1.4%
12
PLOS Computational Biology
1863 papers in training set
Top 16%
1.3%
13
Scientific Reports
3612 papers in training set
Top 66%
1.1%
14
Genetic Epidemiology
55 papers in training set
Top 0.6%
1.1%
15
Genome Biology
637 papers in training set
Top 7%
1.1%
16
Bioinformatics
1204 papers in training set
Top 8%
1.0%
17
Communications Biology
993 papers in training set
Top 27%
1.0%
18
European Journal of Human Genetics
58 papers in training set
Top 1.0%
1.0%
19
PLOS ONE
5266 papers in training set
Top 62%
0.8%
20
Genome Research
468 papers in training set
Top 6%
0.8%
21
Human Molecular Genetics
141 papers in training set
Top 3%
0.8%
22
Briefings in Bioinformatics
354 papers in training set
Top 7%
0.8%
23
Human Genetics
28 papers in training set
Top 0.6%
0.8%