Delphy: scalable, near-real-time Bayesian phylogenetics for outbreaks
Varilly, P.; Schifferli, M.; Yang, K.; Burcham, T.; Cronan, P.; Glennon, O.; Jacks, O.; Laning, E.; Marrs, L.; Oba, K.; Yeung, S.; Parker, E.; Omah, I.; Pekar, J. E.; Luebbert, L.; Andersen, K. G.; Park, D. J.; Schaffner, S. F.; MacInnis, B. L.; Happi, C.; Lemieux, J. E.; Ozonoff, A.; Mitzenmacher, M. D.; Fry, B.; Sabeti, P. C.
Show abstract
Pathogen genomic analysis is central to tracking, understanding, and containing outbreaks, but complexity and high costs of state-of-the-art (SOTA) phylogenetic tools limit global access and impact. We introduce Delphy, an exact reformulation of Bayesian phylogenetics designed to transform its speed, scalability and accessibility while retaining SOTA accuracy. Delphys central data structure, an Explicit Mutation Annotated Tree, exploits the high sequence similarity in large-scale epidemic datasets for efficient tree exploration and convergence. By reproducing key analyses from recent major epidemics (Ebola, Zika, SARS-CoV-2, mpox, and H5N1), we demonstrate SOTA accuracy with up to 1,000x speedups. Assessing Delphys scalability, we show that a simulated dataset of 100,000 sequences can be analyzed in under a day-the largest such computation to date. We distribute Delphy as a client-side web application, enabling users worldwide to turn raw data into interactive results within minutes, without the data ever leaving the users machine. Delphy automatically identifies key viral lineages and mutations, as well as their emergence and prevalence through time, all with quantified uncertainties derived from a solid theoretical foundation. Delphy shows the power of Bayesian phylogenetics as a fast, accessible frontline tool for tackling future outbreaks.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.