The high-throughput gene prediction of more than 1,700 eukaryote genomes using the software package EukMetaSanity
Neely, C. J.; Hu, S. K.; Alexander, H.; Tully, B. J.
Show abstract
Gene prediction and annotation for eukaryotic genomes is challenging with large data demands and complex computational requirements. For most eukaryotes, genomes are recovered from specific target taxa. However, it is now feasible to reconstruct or sequence hundreds of metagenome-assembled genomes (MAGs) or single-amplified genomes directly from the environment. To meet this forth-coming wave of eukaryotic genome generation, we introduce EukMetaSanity, which combines state-of-the-art tools into three pipelines that have been specifically designed for extensive parallelization on high-performance computing infrastructure. EukMetaSanity performs an automated taxonomy search against a protein database of 1,482 species to identify phylogenetically compatible proteins to be used in downstream gene prediction. We present the results for intron, exon, and gene locus prediction for 112 genomes collected from NCBI, including fungi, plants, and animals, along with 1,669 MAGs and demonstrate that EukMetaSanity can provide reliable preliminary gene predictions for a single target taxon or at scale for hundreds of MAGs. EukMetaSanity is freely available at https://github.com/cjneely10/EukMetaSanity.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- mettannotator: a comprehensive and scalable Nextflow annotation pipeline for prokaryotic assemblies 96%
- PhyloMagnet: Fast and accurate screening of short-read meta-omics data using gene-centric phylogenetics 96%
- Evolclust: automated inference of evolutionary conserved gene clusters in eukaryotes 96%
Similar papers in this journal
Similar papers in this journal
- OMAnnotator: a novel approach to building an annotated consensus genome sequence 96%
- MerCat2: a versatile k-mer counter and diversity estimator for database-independent property analysis obtained from omics data 96%
- BugBuster: A novel automatic and reproducible workflow for metagenomic data analysis 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.