Back

EasyPseudogene: an easy-to-use and multithreaded pipeline for pseudogene detection

Ai, C.; Tan, L.; Gao, S.; Wang, Y.

2026-03-06 bioinformatics
10.64898/2026.03.04.709571 bioRxiv
Show abstract

Pseudogenes are recognized as essential components for reconstructing adaptive evolutionary trajectories and understanding genomic remodeling. However, identifying these sequences in large eukaryotic genomes remains technically challenging due to fragmented workflows, complex manual configurations, and the lack of high-performance, parallelized tools capable of processing rapidly growing data volumes. We present EasyPseudogene, an automated and multithreaded pipeline designed for the end-to-end identification of pseudogenes across diverse eukaryotic lineages. Unlike traditional self-mapping tools that often fail to detect unitary pseudogenes when functional counterparts are absent, EasyPseudogene introduces an inter-species reference-driven paradigm that utilizes high-quality proteomes as probes to scan target genomes for evolutionary relics. The pipeline employs a modular "hierarchical screening and precision detection" architecture, integrating high-speed homology searching via MMseqs2 and spliced alignments via miniprot with high-fidelity, three-frame alignments using GeneWise. Performance benchmarking on cetacean genomes demonstrates that EasyPseudogene can replicate known gene loss events, such as the functional decay of the ADRB3 gene, with 100% consistency relative to established manual workflows. By encapsulating complex comparative genomics logic into a standardized framework with interactive HTML visualization for mutation auditing at single-base resolution, EasyPseudogene provides a versatile and reproducible solution for marine ecology and evolutionary research.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.