Accurate de novo identification of biosynthetic gene clusters with GECCO
Carroll, L. M.; Larralde, M.; Fleck, J. S.; Ponnudurai, R.; Milanese, A.; Cappio Barazzone, E.; Zeller, G.
Show abstract
Biosynthetic gene clusters (BGCs) are enticing targets for (meta)genomic mining efforts, as they may encode novel, specialized metabolites with potential uses in medicine and biotechnology. Here, we describe GECCO (GEne Cluster prediction with COnditional random fields; https://gecco.embl.de), a high-precision, scalable method for identifying novel BGCs in (meta)genomic data using conditional random fields (CRFs). Based on an extensive evaluation of de novo BGC prediction, we found GECCO to be more accurate and over 3x faster than a state-of-the-art deep learning approach. When applied to over 12,000 genomes, GECCO identified nearly twice as many BGCs compared to a rule-based approach, while achieving higher accuracy than other machine learning approaches. Introspection of the GECCO CRF revealed that its predictions rely on protein domains with both known and novel associations to secondary metabolism. The method developed here represents a scalable, interpretable machine learning approach, which can identify BGCs de novo with high precision.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- SHARK enables homology assessment in unalignable anddisordered sequences 95%
- Revealing 29 sets of independently modulated genes in Staphylococcus aureus, their regulators and role in key physiological responses 95%
- Versatile NTP recognition and domain fusions expand the functional repertoire of the ParB-CTPase fold beyond chromosome segregation 94%
Similar papers in this journal
- Long-read metagenomics of soil communities reveals phylum-specific secondary metabolite dynamics 94%
- BugSplit: highly accurate taxonomic binning of metagenomic assemblies enables genome-resolved metagenomics 94%
- Unravelling the Core and Accessory Genome Diversity of Enterobacteriaceae in Carbon Metabolism 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.