Back

hybpiper-rbgv and yang-and-smith-rbgv: Containerization and additional options for assembly and paralog detection in target enrichment data

Jackson, C.; McLay, T.; Schmidt-Lebuhn, A. N.

2021-11-10 bioinformatics
10.1101/2021.11.08.467817 bioRxiv
Show abstract

PREMISEThe HybPiper pipeline has become one of the most widely used tools for the assembly of target enrichment (sequence capture) data for phylogenomic analysis. Between the production of locus sequences and phylogenetic analysis, the identification of paralogs is a critical step ensuring accurate inference of evolutionary relationships. Algorithmic approaches using gene tree topologies for the inference of ortholog groups are computationally efficient and broadly applicable to non-model organisms, especially in the absence of a known species tree. Unfortunately, software compatibility issues, unfamiliarity with relevant programming languages, and the complexity involved in running numerous subsequent analysis steps continue to limit the broad uptake of these approaches and constrain their application in practice. METHODS AND RESULTSWe updated the scripts constituting HybPiper and a pipeline for the inference of ortholog groups ("Yang and Smith") to provide novel options for the treatment of supercontigs, remove bugs, and seamlessly use the outputs of the former as inputs for the latter. The pipelines were containerised using Singularity and implemented via two Nextflow pipelines for easier deployment and to vastly reduce the number of commands required for their use. We tested the pipelines with several datasets, one of which is presented for demonstration. CONCLUSIONShybpiper-rbgv and yang-and-smith-rbgv provide easy installation, user-friendly experience, and robust results to the phylogenetic community. They are presently used as the analysis pipeline of the Australian Angiosperm Tree of Life project. The pipelines are available at https://github.com/chrisjackson-pellicle.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
Applications in Plant Sciences
23 papers in training set
Top 0.1%
37.9%
2
Molecular Biology and Evolution
542 papers in training set
Top 1.0%
7.6%
3
Systematic Biology
144 papers in training set
Top 0.3%
6.4%
50% of probability mass above
4
Bioinformatics
1204 papers in training set
Top 4%
5.3%
5
Molecular Ecology Resources
171 papers in training set
Top 0.4%
5.2%
6
Methods in Ecology and Evolution
176 papers in training set
Top 0.5%
5.2%
7
PeerJ
308 papers in training set
Top 2%
4.1%
8
BMC Bioinformatics
457 papers in training set
Top 3%
2.5%
9
Systematic Entomology
14 papers in training set
Top 0.1%
2.0%
10
The Plant Journal
215 papers in training set
Top 3%
1.6%
11
Bioinformatics Advances
203 papers in training set
Top 4%
1.3%
12
G3: Genes, Genomes, Genetics
252 papers in training set
Top 3%
1.3%
13
PLOS Computational Biology
1863 papers in training set
Top 18%
1.1%
14
Frontiers in Plant Science
256 papers in training set
Top 4%
1.1%
15
NAR Genomics and Bioinformatics
242 papers in training set
Top 4%
1.0%
16
Plant Direct
95 papers in training set
Top 2%
1.0%
17
Genome Biology and Evolution
338 papers in training set
Top 3%
0.8%
18
Horticulture Research
47 papers in training set
Top 1.0%
0.8%
19
BMC Genomics
406 papers in training set
Top 9%
0.8%
20
GigaScience
212 papers in training set
Top 5%
0.8%
21
G3 Genes|Genomes|Genetics
351 papers in training set
Top 4%
0.6%
22
Plant Physiology
238 papers in training set
Top 3%
0.6%