PASV: Automatic protein partitioning and validation using conserved residues
Moore, R. M.; Harrison, A. O.; Nasko, D. J.; Chopyk, J.; Cebeci, M.; Ferrell, B. D.; Polson, S. W.; Wommack, K. E.
Show abstract
BackgroundIncreasingly, researchers use protein-coding genes from targeted PCR amplification or direct metagenomic sequencing in community and population ecology. Analysis of protein-coding genes presents different challenges from those encountered in traditional SSU rRNA studies. Most protein-coding sequences are annotated based on homology to other computationally-annotated sequences, which can lead to inaccurate annotations. Therefore, the results of sensitive homology searches must be validated to remove false-positives and assess functionality. Multiple lines of in silico evidence can be gathered by examining conserved domains and residues identified through biochemical investigations. However, manually validating sequences in this way can be time consuming and error prone, especially in large environmental studies. ResultsAn automated pipeline for protein active site validation (PASV) was developed to improve validation and partitioning accuracy for protein-coding sequences, combining multiple sequence alignment with expert domain knowledge. PASV was tested using commonly misannotated proteins: ribonucleotide reductase (RNR), alternative oxidase (AOX), and plastid terminal oxidase (PTOX). PASV partitioned 9,906 putative Class I alpha and Class II RNR sequences from bycatch in a global viral metagenomic investigation with >99% true positive and true negative rates. PASV predicted the class of 2,579 RNR sequences in >98% agreement with manual annotations. PASV correctly partitioned all 336 tested AOX and PTOX sequences. ConclusionsPASV provides an automated and accurate way to address post-homology search validation and partitioning of protein-coding marker genes. Source code is released under the MIT license and is found with documentation and usage examples on GitHub at https://github.com/mooreryan/pasv.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- AutoPhy: Automated phylogenetic identification of novel protein subfamilies 96%
- Comparative evaluation of bioinformatic tools for virus-host prediction and their application to a highly diverse community in the Cuatro Cienegas Basin, Mexico 96%
- Short k-mer Abundance Profiles Yield Robust Machine Learning Features and Accurate Classifiers for RNA Viruses 95%
Similar papers in this journal
Similar papers in this journal
- Ribovore: ribosomal RNA sequence analysis for GenBank submissions and database curation 96%
- Profile hidden Markov model sequence analysis can help remove putative pseudogenes from DNA barcoding and metabarcoding datasets 95%
- VADR: validation and annotation of virus sequence submissions to GenBank 95%
Similar papers in this journal
- Identifying Genomic Islands with Deep Neural Networks 94%
- MeShClust v3.0: High-quality clustering of DNA sequences using the mean shift algorithm and alignment-free identity scores 93%
- Isolation and Characterization of a Roseophage Representing a Novel Genus in the N4-like Rhodovirinae Subfamily Distributed in Estuarine Waters 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.