Sidero-Mining: Systematic Extraction of Siderophore Biosynthetic Information Using Large Language Models
Shao, J.; Wu, Y.; Xi, N.; He, R.; Xu, R.; Guo, P.; Gu, S.; Li, Z.
Show abstract
Siderophores are essential secondary metabolites widely distributed across microorganisms, displaying remarkable diversity. Despite extensive research, public databases contain limited information on siderophore biosynthetic gene clusters (BGCs), particularly lacking cross-species distribution and biosynthetic substrate annotations. Systematically collecting and organizing siderophore BGC synthesis data on a large scale would significantly enhance the use of domain knowledge and support data-driven research. Large language models (LLMs) now offer a practical and scalable approach for mining and curating biological data, especially for converting literature insights into structured datasets In this work, we developed the Sidero-Mining pipeline, using LLMs to efficiently extract siderophore BGC synthesis information. By employing LLMs to screen over 10,000 publications, we identified 1,843 high-quality articles for data mining based on Sidero-Mining framework, manual validation, and data integration. This effort culminated in the creation of the most comprehensive siderophore BGC dataset to date, containing 728 BGCs and 325 NRPS A domain substrate entries cross various species. Our results highlight LLMs potential to accelerate secondary metabolite dataset construction, and our methodological framework can be adapted for systematically exploring other secondary metabolites.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deciphering the Biosynthetic Potential of Microbial Genomes Using a BGC Language Processing Neural Network Model 95%
- PanKB: An interactive microbial pangenome knowledgebase for research, biotechnological innovation, and knowledge mining 95%
- A Deep Learning Genome-Mining Strategy Improves Biosynthetic Gene Cluster Prediction 95%
Similar papers in this journal
- HVRLocator: A Computationally Efficient Tool for Identifying Hypervariable Regions in 16S rRNA Big Datasets 91%
- TooManyCellsInteractive: a visualization tool for dynamic exploration of single-cell data 91%
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database 90%
Similar papers in this journal
- SIDERITE: Unveiling Hidden Siderophore Diversity in the Chemical Space Through Digital Exploration 93%
- ViWrap: A modular pipeline to identify, bin, classify, and predict viral-host relationships for viruses from metagenomes 90%
- High-throughput generic single-entity sequencing using droplet microfluidics 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.