Back

Splicer: Phylogenetic Placement in Sub-Linear Time

Markin, A.; Anderson, T. K.

2026-02-12 bioinformatics
10.64898/2026.02.10.705130 bioRxiv
Show abstract

MotivationPhylogenetic placement is an established approach for rapidly classifying new genetic sequences and updating a phylogeny without fully recomputing it. Popular maximum-likelihood placement methods, such as pplacer and EPA-ng, tend to struggle computationally when the size of the reference tree increases to tens or hundreds of thousands of sequences. As a more scalable alternative, distance-based and parsimony-based placement methods were introduced such as UShER. These methods, in principle, scale linearly as the size of the reference tree grows. However, as the scale of genetic and genomic sequences continues to grow nearly exponentially, developing algorithms that can perform placement in sub-linear time while maintaining accuracy becomes more crucial. ResultsHere, we develop Splicer, the first such algorithm that can perform placement in guaranteed [Formula] time. To achieve this performance, Splicer first decomposes the original reference tree into blobs and constructs a phylogenetic scaffold tree linking representatives from different blobs. Every blob in such decomposition has at most [Formula] taxa, and the scaffold tree has at most [Formula] leaves, where c is any constant. Then, given the query sequences for placement, they are first placed onto a scaffold tree using pplacer or EPA-ng, and then placed more precisely within the respective blobs. We demonstrate the high accuracy of Splicer on an empirical influenza A virus dataset that has sparse coverage due to limited genomic surveillance. We also show that Splicer can, for the first time, apply maximum-likelihood placement to COVID-19 pandemic-scale data using a dataset with over 12 million SARS-CoV-2 reference genomes. Splicer scales the highly accurate maximum-likelihood approaches implemented in pplacer and EPA-ng to trees with millions of taxa and eliminates the necessity to curate and subsample genomic datasets for real-time classifications. Availability and implementationSplicer tool and source code are freely available at https://github.com/flu-crew/splicer.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.