Back

doubletrouble: an R/Bioconductor package for the identification, classification, and analysis of gene and genome duplications

Almeida-Silva, F.; Van de Peer, Y.

2024-02-29 bioinformatics
10.1101/2024.02.27.582236 bioRxiv
Show abstract

Gene and genome duplications are major evolutionary forces that shape the diversity and complexity of life. However, different duplication modes have distinct impacts on gene function, expression, and regulation. Existing tools for identifying and classifying duplicated genes are either outdated or not user-friendly. Here, we present doubletrouble, an R/Bioconductor package that provides a comprehensive and robust framework for analyzing duplicated genes from genomic data. doubletrouble can detect and classify gene pairs as derived from six duplication modes (segmental, tandem, proximal, retrotransposon-derived, DNA transposon-derived, and dispersed duplications), calculate substitution rates, detect signatures of putative whole-genome duplication events, and visualize results as publication-ready figures. We applied doubletrouble to classify the duplicated gene repertoire in 822 eukaryotic genomes, which we made available through a user-friendly web interface (available at https://almeidasilvaf.github.io/doubletroubledb). doubletrouble is freely accessible from Bioconductor (https://bioconductor.org/packages/doubletrouble), and it provides a valuable resource to study the evolutionary consequences of gene and genome duplications.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.