Back

The impact of incomplete taxon sampling on inference of gene flow by Bayesian and summary methods using genomic sequence data

Cheng, S.; Flouri, T.; Zhu, T.; Yang, Z.

2025-11-06 genomics
10.1101/2025.11.05.686701 bioRxiv
Show abstract

Interspecific gene flow is commonly inferred using genomic data under the multispecies coalescent (MSC) model. The rate of gene flow is measured by the expected proportion of immigrants in the recipient population at the time of hybridization/introgression. Incomplete taxon sampling can impact inference of gene flow in multiple ways. First unsampled ghost lineages that are sources of introgression may mislead inference of gene flow in analysis of genomic data from modern species. Second incomplete taxon sampling causes merges of branches on the species phylogeny, which represent populations of different sizes, and complicates the definition and estimation of the introgression probability. We use mathematical analysis and computer simulation to examine the impact of incomplete taxon sampling on inference of gene flow and estimation of its rate using genomic data. We introduce a Bayesian testing approach to select models of gene flow for a species triplet (such as inflow, outflow, and ghost introgression), using the Savage-Dickey density ratio to calculate Bayes factors. We show that the approach has excellent power and specificity. We find that genomic data allow reliable estimation of the proportion of immigrants (rather than the number of immigrants), even when the assumed demographic model is incorrect due to incomplete taxon sampling. When population size differs among species, assuming the same size may lead to seriously biased estimates of the rate of gene flow. The f -branch approach is found to be effective in reducing the number of significant gene-flow events from triplet analyses, providing useful hypotheses for rigorous testing, but often to produce underestimates of the rate of gene flow. Our results highlight the need for improving summary methods to accommodate different population sizes and to infer gene flow between sister lineages.

Published in Systematic Biology (predicted rank #1) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.