BioCAT: PSSM-based algorithm to search biosynthetic gene clusters producing nonribosomal peptides with known structure
Konanov, D. N.; Krivonos, D. V.; Babenko, V. V.; Ilina, E. N.
Show abstract
MotivationNonribosomal peptides are a class of secondary metabolites synthesized by multimodular enzymes named nonribosomal peptide synthetases and mainly produced by bacteria and fungi. It has been shown that non-ribosomal peptides have a huge structural and functional diversity including antimicrobial activity, therefore, they are of increasing interest for modern biotechnology. Methods such as NMR and LC-MS/MS allow to determine a peptide structure precisely, but it is often not a trivial task to find natural producers of them. Today, the search is usually performed manually, mostly with tools such as antiSMASH or Prism. However, there are cases when potential producers should be found among hundreds of strains, for instance, when analyzing metagenomes data. Thus, the development of automated approaches is a high-priority task for further nonribosomal peptides research. ResultsWe developed BioCAT, a two-side approach to find biosynthetic gene clusters which may produce a given nonribosomal peptide when the structure of interesting nonribosomal peptide has already been found. Formally, BioCAT unites the antiSMASH software and the rBAN retrosynthesis tool but some improvements were added to both gene cluster and peptide chemical structure analyses. The main feature of the method is an implementation of position specific score matrix to store specificities of nonribosomal peptide synthetase modules, which has increased the alignment quality in comparison with more strict approaches developed earlier. An ensemble model was implemented to calculate the final alignment score. We tested the method on a manually curated nonribosomal peptides producers database and compared it with a competing tool called GARLIC. Finally, we showed the method applicability on several external examples. AvailabilityBioCAT is available on the GitHub repository or via pip Contactkonanovdmitriy@gmail.com
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Predicting biological pathways of chemical compounds with a profile-inspired aproach 94%
- Predicting the pathway involvement of metabolites annotated in the MetaCyc knowledgebase 94%
- Binding affinity prediction for protein-ligand complex using deep attention mechanism based on intermolecular interactions 94%
Similar papers in this journal
Similar papers in this journal
- A Novel Riboswitch Classification based on Imbalanced Sequences achieved by Machine Learning 94%
- Elucidation of Genome-wide Understudied Proteins targeted by PROTAC-induced degradation using Interpretable Machine Learning 94%
- Integrating structure-based machine learning and co-evolution to investigate specificity in plant sesquiterpene synthases 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.