Read2Tree: scalable and accurate phylogenetic trees from raw reads
Dylus, D.; Altenhoff, A. M.; Majidian, S.; Sedlazeck, F. J.; Dessimoz, C.
Show abstract
The inference of phylogenetic trees is foundational to biology. However, state-of-the-art phylogenomics requires running complex pipelines, at significant computational and labour costs, with additional constraints in sequencing coverage, assembly and annotation quality. To overcome these challenges, we present Read2Tree, which directly processes raw sequencing reads into groups of corresponding genes. In a benchmark encompassing a broad variety of datasets, our assembly-free approach was 10-100x faster than conventional approaches, and in most cases more accurate--the exception being when sequencing coverage was high and reference species very distant. To illustrate the broad applicability of the tool, we reconstructed a yeast tree of life of 435 species spanning 590 million years of evolution. Applied to Coronaviridae samples, Read2Tree accurately classified highly diverse animal samples and near-identical SARS-CoV-2 sequences on a single tree--thereby exhibiting remarkable breadth and depth. The speed, accuracy, and versatility of Read2Tree enables comparative genomics at scale.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Purging genomes of contamination eliminates systematic bias from evolutionary analyses of ancestral genomes 96%
- Benchmarking strategies for cross-species integration of single-cell RNA sequencing data 95%
- A comparative analysis of planarian genomes reveals regulatory conservation in the face of rapid structural divergence 95%
Similar papers in this journal
Similar papers in this journal
- vRhyme enables binning of viral genomes from metagenomes 96%
- NeMu: A Comprehensive Pipeline for Accurate Reconstruction of Neutral Mutation Spectra from Evolutionary Data 95%
- Domainator, a flexible software suite for domain-based annotation and neighborhood analysis, identifies proteins involved in antiviral systems 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.