OrthoPhyl - A turn-key solution for large scale whole genome bacterial phylogenomics
Middlebrook, E.; Katani, R.; Fair, J. M.
Show abstract
There are a staggering number of publicly available bacterial genome sequences (at writing, 2.0 million assemblies in NCBIs GenBank alone), and the deposition rate continues to increase. This wealth of data begs for phylogenetic analyses to place these sequences within an evolutionary context. A phylogenetic placement not only aids in taxonomic classification, but informs the evolution of novel phenotypes, targets of selection, and horizontal gene transfer. Building trees from multi-gene codon alignments is a laborious task that requires bioinformatic expertise, rigorous curation of orthologs, and heavy computation. Compounding the problem is the lack of tools that can streamline these processes for building trees from large scale genomic data. Here we present OrthoPhyl, which takes bacterial genome assemblies and reconstructs trees from whole genome codon alignments. The analysis pipeline can analyze an arbitrarily large number of input genomes (>1200 tested here) by identifying a diversity spanning subset of assemblies and using these genomes to build gene models to infer orthologs in the full dataset. To illustrate the versatility of OrthoPhyl, we show three use-cases: E. coli/Shigella, Brucella/Ochrobactrum, and the order Rickettsiales. We compare trees generated with OrthoPhyl to trees generated with kSNP3 and GToTree along with published trees using alternative methods. We show that OrthoPhyl trees are consistent with other methods while incorporating more data, allowing for greater numbers of input genomes, and more flexibility of analysis. Availability and ImplementationCode used in this manuscript is available at https://github.com/eamiddlebrook/OrthoPhyl/blob/OrthoPhyl_1.0/. Installation and execution instructions are provided in the associated github README.md file. Third party software versions and OrthoPhyl execution files will remain static in the OrthoPhyl_1.0 branch, with the main branch housing current development. For versions of software dependencies see Supplemental Table 1. To aid in usability, a Singularity container is available at https://cloud.sylabs.io/library/earlyevol/default/orthophyl or with the command singularity pull library://earlyevol/default/orthophyl:1.0_ms. Also see the GitHub page for Singularity usage guide. Assemblies used within this manuscript are available from https://www.ncbi.nlm.nih.gov/assembly/ with accessions in Supplemental tables 2-4.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- PhyloMagnet: Fast and accurate screening of short-read meta-omics data using gene-centric phylogenetics 96%
- HIPSTR: highest independent posterior subtree reconstruction in TreeAnnotator X 95%
- TopHap: Rapid inference of key phylogenetic structures from common haplotypes in large genome collections with limited diversity 94%
Similar papers in this journal
- Rephine.r: a pipeline for correcting gene calls and clusters to improve phage pangenomes and phylogenies 95%
- Robust genome-based delineation of bacterial genera 95%
- Mitochondrial genomes of Columbicola feather lice are highly fragmented, indicating repeated evolution of minicircle-type genomes in parasitic lice 93%
Similar papers in this journal
- Broccoli: combining phylogenetic and network analyses for orthology assignment 95%
- PhyloAln: a convenient reference-based tool to align sequences and high-throughput reads for phylogeny and evolution in the omic era 95%
- A daily-updated database and tools for comprehensive SARS-CoV-2 mutation-annotated trees 95%
Similar papers in this journal
- ClusTRace, a bioinformatic pipeline for analyzing clusters in virus phylogenies 95%
- The advantages and disadvantages of short- and long-read metagenomics to infer bacterial and eukaryotic community composition 94%
- cognac: rapid generation of concatenated gene alignments for phylogenetic inferencefrom large whole genome sequencing datasets 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.