Back

Efficient Phylogenetic Inference Using SNP-Based Approaches: A Comparison with Full Sequence Data

Vivekanantha, S.; Konig, P.

2025-03-13 bioinformatics
10.1101/2025.03.10.642383 bioRxiv
Show abstract

Mutations in specific genomic regions or genes serve as reliable indicators of phylogenetic relationships, with single nucleotide polymorphisms (SNPs) playing a crucial role in population phylogenetic studies. Traditional distance-based phylogenetic algorithms have a time complexity proportional to l n2, where n is the number of sequences and l is their length [1], [2]. This high computational cost becomes a bottleneck in phylogeny reconstruction, particularly when l > n. To overcome this limitation, we propose an SNP-based approach to phylogenetic tree inference, focusing exclusively on variant (mutated) positions rather than entire sequences. This method significantly reduces computational time while maintaining accuracy. We compare phylogenies inferred from SNP data and full sequence data (including both SNPs and invariant sites) across multiple metrics. Our results show that heuristic phylogenetic trees constructed from SNPs achieve parsimony scores nearly identical to those derived from full sequence data, as parsimony primarily depends on variant positions. Additionally, under the Jukes-Cantor 1969 (JC69) model, log-likelihood scores for SNP-based and full-sequence-based trees exhibit a strong correlation when evaluated using the same tree topology, branch lengths, and maximum likelihood parameters. These findings demonstrate that SNP-based methods can streamline phylogenetic analysis while preserving accuracy.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.