Back

Optimal Reference Panel Design in Ancient DNA Imputation from Coalescent Theory, Simulation, and Real Data Application with an Ancient Reference Panel

Sousa da Mota, B.; Kumar, K.; Reich, D. E.; Zoellner, S.

2026-04-28 genomics
10.64898/2026.04.27.721163 bioRxiv
Show abstract

Imputation is widely used in the ancient DNA (aDNA) field to determine which phenotypically important alleles ancient individuals carried, to study natural selection, and to detect segments of the genome that are shared between individuals identical by descent. However, rare variant imputation is less accurate, and rare variants tend to be excluded from downstream analyses. State-of-the-art imputation methods leverage large reference panels, improving rare variant accuracy in modern targets. However, it is unclear how to identify optimal panels for aDNA targets. It seems plausible that aDNA reference panels would improve imputation of aDNA, but no such panels have been assembled or tested. We leveraged analytical results from coalescent theory and complementary simulations to evaluate both performance of large modern panels, and ancient panels impact on aDNA imputation. For modern panels, sample sizes as small as 5,000 saturate imputation performance and model misspecifications in standard imputation algorithms increase imputation error for rare and intermediate frequency variants. For instance, for European hunter-gatherers, non-reference imputed variants with derived allele frequency less than at least 2% should be removed. Including ancient genomes in a modern reference panel substantially improved imputation accuracy in analytical modelling and simulations, particularly, for rare variants and older samples from groups with low effective population size. We assembled a joint reference panel with 1000 Genomes and 95 ancient samples and used it to impute 95 downsampled genomes, finding modest gains in imputation performance. This approach can rescue rare variants typically discarded from current imputation pipelines and may prove useful as the number of ancient samples increases.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Molecular Ecology Resources
171 papers in training set
Top 0.1%
12.6%
2
Molecular Biology and Evolution
542 papers in training set
Top 0.5%
12.4%
3
Genome Biology
637 papers in training set
Top 1%
9.5%
4
Genome Biology and Evolution
338 papers in training set
Top 0.8%
6.1%
5
The American Journal of Human Genetics
234 papers in training set
Top 0.9%
5.0%
6
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 11%
4.7%
50% of probability mass above
7
Genome Research
468 papers in training set
Top 2%
4.2%
8
BMC Genomics
406 papers in training set
Top 2%
4.0%
9
GENETICS
483 papers in training set
Top 1%
3.4%
10
Nature Communications
5641 papers in training set
Top 36%
3.2%
11
eLife
5828 papers in training set
Top 37%
3.0%
12
Science
477 papers in training set
Top 3%
2.6%
13
PLOS Genetics
862 papers in training set
Top 5%
2.4%
14
Human Genetics and Genomics Advances
84 papers in training set
Top 1.0%
2.1%
15
Journal of Heredity
42 papers in training set
Top 0.3%
2.1%
16
G3: Genes, Genomes, Genetics
252 papers in training set
Top 3%
1.7%
17
Scientific Reports
3612 papers in training set
Top 56%
1.7%
18
PLOS Computational Biology
1863 papers in training set
Top 16%
1.4%
19
Nature
645 papers in training set
Top 8%
1.3%
20
Frontiers in Genetics
230 papers in training set
Top 5%
1.0%
21
Bioinformatics
1204 papers in training set
Top 8%
1.0%
22
Philosophical Transactions of the Royal Society B: Biological Sciences
72 papers in training set
Top 2%
0.8%
23
Nature Genetics
286 papers in training set
Top 5%
0.8%
24
Computational and Structural Biotechnology Journal
242 papers in training set
Top 8%
0.6%
25
Molecular Ecology
336 papers in training set
Top 4%
0.6%
26
Nucleic Acids Research
1281 papers in training set
Top 15%
0.6%
27
GigaScience
212 papers in training set
Top 5%
0.6%
28
Current Biology
665 papers in training set
Top 11%
0.6%
29
PeerJ
308 papers in training set
Top 13%
0.6%