Back

On the impact of reference selection on variant calling in phylogenomics: How to avoid systematic error in target-enrichment studies of non-model organisms

Zhang, G.; Koehler, F.

2025-03-25 evolutionary biology
10.1101/2025.03.23.644254 bioRxiv
Show abstract

Variant calling is a crucial step in identifying genetic variation in DNA sequences reconstructed from assembled sequencer reads. This study investigates the effect of reference genome selection on the reliability of variant calling and assesses the consequences of reference genome choice on phylogenomic analyses downstream. Employing an empirical exon capture sequence dataset, we used an analysis pipeline implemented in the Reference Genome based Phylogeny Pipeline (RGBEPP) that includes steps of quality control, read mapping, variant calling, and phylogenetic analysis. To examine how variant detection is influenced by reference genome choice, we assembled the exon capture dataset by using four different references for variant calling. These references were (1) a single ingroup sample, (2) the combined ingroup samples, (3) a distantly related species, and (4) a self-derived reference for each sample. We found that reference choice significantly impacted the variant detection and that these differences in variant detection also influenced the phylogenetic reconstructions downstream. Comparing the alignments and trees produced under use of various references shows that employing a sample-specific self-reference for variant calling maximizes the accuracy of the variant detection process. Based on this finding, we recommend incorporating self-referenced variant calling into phylogenomic assembly and analysis pipelines, such as RGBEPP, to ensure the robustness and reproducibility of phylogenomic analyses of low coverage sequence datasets from non-human species.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.