doubletrouble: an R/Bioconductor package for the identification, classification, and analysis of gene and genome duplications
Almeida-Silva, F.; Van de Peer, Y.
Show abstract
Gene and genome duplications are major evolutionary forces that shape the diversity and complexity of life. However, different duplication modes have distinct impacts on gene function, expression, and regulation. Existing tools for identifying and classifying duplicated genes are either outdated or not user-friendly. Here, we present doubletrouble, an R/Bioconductor package that provides a comprehensive and robust framework for analyzing duplicated genes from genomic data. doubletrouble can detect and classify gene pairs as derived from six duplication modes (segmental, tandem, proximal, retrotransposon-derived, DNA transposon-derived, and dispersed duplications), calculate substitution rates, detect signatures of putative whole-genome duplication events, and visualize results as publication-ready figures. We applied doubletrouble to classify the duplicated gene repertoire in 822 eukaryotic genomes, which we made available through a user-friendly web interface (available at https://almeidasilvaf.github.io/doubletroubledb). doubletrouble is freely accessible from Bioconductor (https://bioconductor.org/packages/doubletrouble), and it provides a valuable resource to study the evolutionary consequences of gene and genome duplications.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Accurate reconstruction of bacterial pan- and core- genomes with PEPPAN 95%
- HiCanu: accurate assembly of segmental duplications, satellites, and allelic variants from high-fidelity long reads 94%
- BRAKER3: Fully Automated Genome Annotation Using RNA-Seq and Protein Evidence with GeneMark-ETP, AUGUSTUS and TSEBRA 94%
Similar papers in this journal
- EASYstrata: An All-in-One Workflow for Genome Annotation and Genomic Divergence Analysis 96%
- PyOrthoANI, PyFastANI, and Pyskani: a suite of Python libraries for computation of average nucleotide identity 95%
- iLoci: Robust evaluation of genome content and organization for provisional and mature genome assemblies 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.