QuickProt: A Fast and Accurate Homology-Based Protein Annotation Tool for Non-Model Organism Genomes
Chen, G.; Du, H.; Cao, Z.; Wu, Y.; Zhang, C.; Zhou, Y.; Ao, J.; Sun, Y.; Yuan, Z.
Show abstract
Despite the increasing number of genomes, the annotation of protein-coding genes is still lacking. In some genomes where transcriptome data is not available, various tools have been developed to annotate genes based on homologous proteins, but each tool has its limitations. Here, we propose quickprot, a user-friendly command-line tool that constructs a set of redundancy-free gene models from the target genome based on the homologous protein sequences of a set of closely related species. We tested its applicability in stonefish, and the results showed that it has high accuracy and runs faster than existing tools. This low-cost, high-precision annotation method is expected to become a useful tool for genome annotation, which will contribute to the research of comparative genomics. The software is implemented in Python and licensed under the MIT License. The source code, documentation, and tutorials can be obtained at https://github.com/thecgs/quickprot.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- TREEasy: an automated workflow to infer gene trees, species trees, and phylogenetic networks from multilocus data 96%
- Chromosome-level genome assembly of the greenfin horse-faced filefish (Thamnaconus septentrionalis) using Oxford Nanopore PromethION sequencing and Hi-C technology 96%
- NGSEP 4: Efficient and Accurate Identification of Orthogroups and Whole Genome Alignment 95%
Similar papers in this journal
- Non-synonymous to synonymous substitutions suggest that orthologs tend to keep their functions, while paralogs are a source of functional novelty 94%
- NGScloud2: optimized bioinformatic analysis using Amazon Web Services 93%
- Genomic comparison of non-photosynthetic plants from the family Balanophoraceae with their photosynthetic relatives. 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.