Variational Inference with Node Embeddings (VINE) for Scalable Bayesian Phylogenetics
Siepel, A.; Hassett, R.; Staklinski, S. J.
Show abstract
Bayesian phylogenetic inference is now widely used but remains heavily reliant on Markov chain Monte Carlo (MCMC) sampling, which is computationally intensive and requires careful convergence monitoring. Variational inference (VI) is an appealing alternative that approximates posterior distributions without sampling, but existing variational approaches for phylogenetics have seen limited adoption owing to constraints in accuracy, speed, and scalability. Here we introduce Variational Inference with Node Embeddings (VO_SCPLOWINEC_SCPLOW), a variational phylogenetic inference method with striking improvements over prior work. VO_SCPLOWINEC_SCPLOW supports both standard DNA substitution models and CRISPR barcode-mutation models for cell-lineage phylogenies. Its key innovations are: embedding taxa in a high-dimensional Euclidean space; backpropagating gradients through fast distance-based phylogeny inference algorithms; introducing a sampling-free approximate estimator for the VI evidence lower bound; and enhancing posterior flexibility using normalizing flows. Across simulated and empirical datasets, VO_SCPLOWINEC_SCPLOW yields accurate posterior approximations for datasets with as many as 1000 taxa in a fraction of the time required for state-of-the-art MCMC-based methods.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DEPP: Deep Learning Enables Extending Species Trees using Single Genes 98%
- ConvexML: Fast and accurate branch length estimation under irreversible mutation models, illustrated through applications to CRISPR/Cas9-based lineage tracing 97%
- Ambiguity coding allows accurate inference of evolutionary parameters from alignments in an aggregated state-space 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.