EasyPseudogene: an easy-to-use and multithreaded pipeline for pseudogene detection
Ai, C.; Tan, L.; Gao, S.; Wang, Y.
Show abstract
Pseudogenes are recognized as essential components for reconstructing adaptive evolutionary trajectories and understanding genomic remodeling. However, identifying these sequences in large eukaryotic genomes remains technically challenging due to fragmented workflows, complex manual configurations, and the lack of high-performance, parallelized tools capable of processing rapidly growing data volumes. We present EasyPseudogene, an automated and multithreaded pipeline designed for the end-to-end identification of pseudogenes across diverse eukaryotic lineages. Unlike traditional self-mapping tools that often fail to detect unitary pseudogenes when functional counterparts are absent, EasyPseudogene introduces an inter-species reference-driven paradigm that utilizes high-quality proteomes as probes to scan target genomes for evolutionary relics. The pipeline employs a modular "hierarchical screening and precision detection" architecture, integrating high-speed homology searching via MMseqs2 and spliced alignments via miniprot with high-fidelity, three-frame alignments using GeneWise. Performance benchmarking on cetacean genomes demonstrates that EasyPseudogene can replicate known gene loss events, such as the functional decay of the ADRB3 gene, with 100% consistency relative to established manual workflows. By encapsulating complex comparative genomics logic into a standardized framework with interactive HTML visualization for mutation auditing at single-base resolution, EasyPseudogene provides a versatile and reproducible solution for marine ecology and evolutionary research.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- PseudoChecker2 and PseudoViz: automation and visualization of gene loss in the Genome Era 96%
- OMAnnotator: a novel approach to building an annotated consensus genome sequence 95%
- MerCat2: a versatile k-mer counter and diversity estimator for database-independent property analysis obtained from omics data 95%
Similar papers in this journal
Similar papers in this journal
- PyOrthoANI, PyFastANI, and Pyskani: a suite of Python libraries for computation of average nucleotide identity 95%
- EASYstrata: An All-in-One Workflow for Genome Annotation and Genomic Divergence Analysis 94%
- iLoci: Robust evaluation of genome content and organization for provisional and mature genome assemblies 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.