Back

Hydroplane I: one-shot probabilistic evolutionary analysis for scalable organizational identification.

Johnson, C. G.; Pollock, D. D.

2025-10-15 evolutionary biology
10.1101/2025.10.14.682387 bioRxiv
Show abstract

Applying evolutionary genomics to microbial and viral community sequence information presents significant challenges. Metagenomic sequences are typically stored in large databases of short-read fragments with unknown relationships. Available tools for analysis are often slow, rely on incomplete subsets of data, or focus on narrowly defined sub-problems. Moreover, existing tools often depend on simplistic model assumptions or treat inferences as empirical data, which can distort downstream analyses. In this work, we present theory and initial validation of the first stage of a fast, one-shot Bayesian approach to metagenomic evolutionary analysis. Our primary result here demonstrates effective use of simple models in experimental design using collections of short proximal short oligonucleotide sequences (kmers) to detect probable homologs, and validates the speed, sensitivity and specificity of this approach in a test set of whole Prochlorococcus genome sequences. Significance StatementThe Hydroplane algorithm provides an efficient framework for analyzing co-evolutionary relationships among microbial and viral genomes using metagenomic data. Hydroplane guides the identification of homologous regions, estimates lineage diversity, and reveals evolutionary events such as mutation, divergence, selection, and horizontal gene transfer. Its computationally lightweight design supports scalable genome clustering and adaptation analyses, enabling the study of rare microbial and viral genomes across diverse ecological contexts. With an accessible implementation, Hydroplane broadens the scope of evolutionary genomics research, offering critical insights into host-pathogen dynamics and supporting biopreparedness through the study of microbial and viral genetic relationships.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.