Back

Simulating population pangenomes under coalescent demographic models with MSpangenome

Piat, L.; Denni, S.; Dubois, S.; Linard, B.; Duvaux, L.

2026-07-03 bioinformatics
10.64898/2026.06.29.735168 bioRxiv
Show abstract

Motivation: Pangenome variation graphs (PVGs) are increasingly used to represent genomic diversity, yet there is currently no general framework for generating population pangenomes directly from explicit evolutionary histories. Existing simulators typically focus on individual classes of variation and do not integrate these variations within a genealogy-aware framework driven by explicit demographic histories. As a result, evaluating pangenome methods in realistic population-genetic settings remains challenging, and benchmark datasets with known evolutionary ground truth are scarce. Results: We present MSpangenome, a genealogy-aware frame- work that bridges coalescent population genetic simulations and pangenome graph analyses. The pipeline combines ancestry simulation with msprime and a de novo graph construction algorithm to generate PVGs directly from simulated genealogies. By explicitly modeling recombination, demographic history and incomplete lineage sorting, MSpangenome produces structurally complex pangenomes in which nested and overlapping structural variants emerge naturally from the underlying genealogies, while their evolutionary history and graph topology remain known by construction. This provides a general framework for generating realistic population pangenomes and establishing ground-truth datasets for methodological evaluation. We demonstrate its utility by generating population-scale pangenomes and using them as controlled references to benchmark the widely used graph construction tools, PGGB and Minigraph-Cactus. Our analyses reveal contrasting performance regimes across levels of sequence diversity, sample sizes and classes of structural variation, highlighting the value of simulation-based benchmarking for identifying reconstruction errors that are hard to detect using empirical datasets alone. Availability and implementation: MSpangenome is imple- mented in Python, fully containerized, freely available at https://forge.inrae.fr/pangepop/MSpangepop and mirrored at https://github.com/inrae/MSpangepop.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 1%
21.8%
2
Molecular Ecology Resources
171 papers in training set
Top 0.2%
9.4%
3
Genome Biology
637 papers in training set
Top 1%
7.7%
4
Molecular Biology and Evolution
542 papers in training set
Top 1%
6.1%
5
BMC Bioinformatics
457 papers in training set
Top 2%
5.4%
50% of probability mass above
6
Methods in Ecology and Evolution
176 papers in training set
Top 0.5%
5.3%
7
Briefings in Bioinformatics
354 papers in training set
Top 2%
4.2%
8
Bioinformatics Advances
203 papers in training set
Top 1%
4.2%
9
PLOS Computational Biology
1863 papers in training set
Top 9%
3.9%
10
GENETICS
483 papers in training set
Top 2%
3.1%
11
Nucleic Acids Research
1281 papers in training set
Top 7%
2.3%
12
NAR Genomics and Bioinformatics
242 papers in training set
Top 2%
2.1%
13
GigaScience
212 papers in training set
Top 2%
2.1%
14
Peer Community Journal
281 papers in training set
Top 3%
1.7%
15
Genome Research
468 papers in training set
Top 4%
1.6%
16
eLife
5828 papers in training set
Top 53%
1.5%
17
BMC Genomics
406 papers in training set
Top 5%
1.5%
18
Genome Biology and Evolution
338 papers in training set
Top 2%
1.4%
19
Nature Communications
5641 papers in training set
Top 48%
1.4%
20
Journal of Open Source Software
25 papers in training set
Top 0.3%
1.1%
21
Nature Computational Science
55 papers in training set
Top 1%
1.1%
22
Computational and Structural Biotechnology Journal
242 papers in training set
Top 6%
1.0%
23
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 43%
0.8%
24
The American Journal of Human Genetics
234 papers in training set
Top 3%
0.8%
25
Cell Reports Methods
165 papers in training set
Top 4%
0.8%
26
PLOS Genetics
862 papers in training set
Top 14%
0.6%