Deciphering Biosynthetic Gene Clusters with a Context-aware Protein Language Model
Kang, Z.; Zhang, H.; Liang, C.; Yang, R.; Ye, Y.; Bai, H.; Zhang, Y.; Ning, K.
Show abstract
Microbial secondary metabolites, synthesized by biosynthetic gene clusters (BGCs), offer vast potential for biotechnological applications. Among BGC profiling techniques, computational detection methods face challenges, including time-consuming alignment and reliance on predefined profiles. To address these, we present BGC-Finder, an end-to-end pipeline utilizing protein language models for BGC detection and annotation from microbial genomes and metagenomes. This approach achieves remarkable increase in profiling speed of up to 100-fold, and employs genomic context-aware modeling to facilitate interpretable genetic essentiality assessment and large-scale BGC clustering. BGC-Finder outperformed traditional methods, successfully detecting 9.49% more biosynthetic-core genes and 27.70% more cytochrome P450s in 742 experimentally-validated BGCs. Notably, it retrieved 31 remote biosynthetic homologs from 210 polar marine metagenomes and identified 4,585 BGCs with 6,388 core genes from 256 fungal genomes. These findings highlight BGC-Finders capability to illuminate "microbial biosynthesis dark matter" (sequence-unrelated, function-similar biosynthetic enzymes) and expedite natural product discovery. HighlightsO_LIBGC-Finder is an accurate and ultrafast pipeline leveraging protein language models (pLMs) to predict and annotate biosynthetic gene clusters (BGCs) from microbial genomes and metagenomes. C_LIO_LIThe genomic context-aware model enables interpretable analysis: attention-driven identification of essential biosynthetic genes and embedding-guided BGC clustering. C_LIO_LIBGC-Finder sensitively retrieves remote homologous BGCs from both bacteria and fungi genomes, uncovering hidden microbial biosynthesis dark matter. C_LIO_LIWe discovered a non-ribosomal peptide synthetase (NRPS) family, which involved into function-specific BGCs in two evolutionarily distant fungi. C_LI
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Microbial general model: Leveraging large language model for contextualized microbiome analysis 95%
- An explainable graph neural framework to identify cancer-associated intratumoral microbial communities 93%
- Long-read sequencing reveals extensive DNA methylations in human gut phagenome contributed by prevalently phage-encoded methyltransferases 93%
Similar papers in this journal
- Deciphering the Biosynthetic Potential of Microbial Genomes Using a BGC Language Processing Neural Network Model 98%
- Targeted genome mining with GATOR-GC maps the evolutionary landscape of biosynthetic diversity 95%
- PanKB: An interactive microbial pangenome knowledgebase for research, biotechnological innovation, and knowledge mining 95%
Similar papers in this journal
Similar papers in this journal
- Biosynthetic gene cluster profiling predicts the positive association between antagonism and phylogeny in Bacillus 94%
- Integration of taxonomic signals from MAGs and contigs improves read annotation and taxonomic profiling of metagenomes 94%
- Detecting and phasing minor single-nucleotide variants from long-read sequencing data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.