VEHoP: A Versatile, Easy-to-use, and Homology-based Phylogenomic pipeline accommodating diverse sequences
Li, Y.; Liu, X.; Chen, C.; Qiu, J.-W.; Kocot, K.; Sun, J.
Show abstract
Phylogenomics has emerged as a transformative approach in systematics, conservation biology, and biomedicine, enabling the inference of evolutionary relationships by leveraging hundreds to thousands of genes from genomic or transcriptomic data. However, acquiring high-quality genomes and transcriptomes necessitates samples with intact DNA and RNA, substantial sequencing investments, and extensive bioinformatic processing, such as genome/transcriptome assembly and annotation. This challenge is particularly pronounced for rare or difficult-to-collect species, such as those inhabiting the deep sea, where only fragmented DNA reads are often available due to environmental degradation or suboptimal preservation conditions. To address these limitations, we introduce VEHoP (Versatile, Easy-to-use Homology-based Phylogenomic pipeline), a tool designed to infer protein-coding regions from diverse inputs, including raw reads (short and long), draft genomes, transcriptomes, and annotated genomes. VEHoP automates the generation of orthologous sequence alignments, concatenated matrices, and phylogenetic trees, streamlining phylogenomic analyses for researchers across disciplines. The tool aims to (1) expand taxonomic sampling by accommodating a wide range of input data types and (2) simplify phylogenomic workflows, making them accessible to researchers with varying levels of bioinformatic expertise. We evaluated VEHoPs performance using datasets from oysters, catfish, and insects, demonstrating its ability to produce robust phylogenetic trees with strong bootstrap support, outperforming assembly-free methods. Additionally, we applied VEHoP to reconstruct the phylogeny of the enigmatic deep-sea gastropod order Neomphalida, successfully resolving a well-supported phylogenetic backbone for this poorly understood group. VEHoP is freely available on GitHub (https://github.com/ylify/VEHoP), with dependencies easily installable via Bioconda.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- TREEasy: an automated workflow to infer gene trees, species trees, and phylogenetic networks from multilocus data 96%
- A snakemake toolkit for the batch assembly, annotation, and phylogenetic analysis of mitochondrial genomes and ribosomal genes from genome skims of museum collections. 95%
- NGSEP 4: Efficient and Accurate Identification of Orthogroups and Whole Genome Alignment 95%
Similar papers in this journal
Similar papers in this journal
- PhyloAln: a convenient reference-based tool to align sequences and high-throughput reads for phylogeny and evolution in the omic era 98%
- MitoNGS: an online platform to analyze fish metabarcoding data in high-resolution 96%
- Broccoli: combining phylogenetic and network analyses for orthology assignment 94%
Similar papers in this journal
- Mitochondrial genomes of Columbicola feather lice are highly fragmented, indicating repeated evolution of minicircle-type genomes in parasitic lice 94%
- Robust genome-based delineation of bacterial genera 93%
- The complete mitochondrial genome sequence of Oryctes rhinoceros (Coleoptera: Scarabaeidae) based on long-read nanopore sequencing 93%
Similar papers in this journal
- A transcriptome-based phylogeny of Scarabaeoidea confirms the sister group relationship of dung beetles and phytophagous pleurostict scarabs (Coleoptera) 92%
- Image-based taxonomic classification of bulk biodiversity samples using deep learning and domain adaptation 92%
- Phylogenomic Analysis of Protein-Coding Genes Resolves Complex Gall Wasp Relationships 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.