Back

GHIST 2024: The 1st Genomic History Inference Strategies Tournament

Struck, T. J.; Vaughn, A. H.; Daigle, A.; Ray, D. D.; Noskova, E.; Sequeira, J. J.; Antonets, S.; Alekseevskaya, E.; Grigoreva, E.; Raines, E.; McMaster, E. S.; Kovacs, T. G. L.; Ragsdale, A. P.; Moreno-Estrada, A.; Lotterhos, K. E.; Siepel, A.; Gutenkunst, R. N.

2025-08-11 evolutionary biology
10.1101/2025.08.05.668560 bioRxiv
Show abstract

Evaluating population genetic inference methods is challenging due to the complexity of evolutionary histories, potential model misspecification, and unconscious biases in self-assessment. The Genomic History Inference Strategies Tournament (GHIST) is a community-driven competition designed to evaluate methods for inferring evolutionary history from population genomic data. The inaugural GHIST competition ran from July to November 2024 and featured four demographic history inference challenges of varying complexity: a bottleneck model, a split with isolation model, a secondary contact model with demographic complexity, and an archaic admixture model. Data were provided as error-free VCF files, and participants submitted numerical parameter estimates that were scored by relative root mean squared error. Approximately 60 participants competed, using diverse approaches. Results revealed the current dominance of methods based on site frequency spectra, while highlighting the advantages of flexible model-building approaches for complex demographic histories. We discuss insights regarding the competition and outline the next iteration, which is ongoing with expanded challenge diversity. By providing standardized benchmarks and highlighting areas for improvement, GHIST represents a substantial step toward more reliable inference of evolutionary history from genomic data.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.