Back

DupyliCate - mining, classifying, and characterizing gene duplications

Natarajan, S.; Pucker, B.

2025-10-13 bioinformatics
10.1101/2025.10.10.681656 bioRxiv
Show abstract

Paralogs, copies of a gene, form an important basis for novelty during evolution. Analysis of such gene duplications is important to understand the emergence of novel traits during evolution. DupyliCate is a Python tool that has been developed for this purpose. With the ability to process multiple datasets concurrently, flexible features, and parameters to set species-specific thresholds, DupyliCate offers a high-throughput method for gene copy identification and analysis. The different available parameters and modes are explored in detail based on the Arabidopsis thaliana datasets. Proof of concept for the tool is presented by characterizing well known duplications in different plants, and its broad applicability is demonstrated by running it on diverse datasets including complex plant genome sequences with high heterozygosity. Further, two case studies involving the evolution of flavonol synthase (FLS) genes in Brassicales, and the evolution of flavonol synthesis regulating myeloblastosis (MYB) transcription factors- MYB12 and MYB111 across a large number of plant species, are presented as exemplar use cases. The tools applicability beyond plants is demonstrated on Escherichia coli, Saccharomyces cerevisiae, and Caenorhabditis elegans datasets. DupyliCate is available at: https://github.com/ShakNat/DupyliCate.

Published in Scientific Reports (predicted rank #3) · training set

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.