SynBioGPT: A Retrieval-Augmented Large Language Model Platform for AI-Guided Microbial Strain Development
Mao, Z.; Du, J.; Wang, R.; Li, H.; Guan, J.; Shi, Z.; Liao, X.; Ma, H.
Show abstract
Synthetic biology seeks to engineer microbial cell factories for sustainable bioproduction, yet the optimization of these systems is impeded by the complexity of metabolic engineering and the protracted timelines of iterative design-build-test-learn (DBTL) cycles. Traditional computational approaches, such as constraint-based modeling, provide valuable insights but demand extensive manual curation. Large language models (LLMs) hold promises for automating knowledge extraction and strain design, yet conventional models like GPT-4 suffer from outdated corpora and hallucination errors in domain-specific tasks. SynBioGPT v1.0, a Retrieval-Augmented Generation (RAG)-enhanced LLM, enhanced knowledge retrieval using vector search but often retrieved semantically similar yet contextually irrelevant documents. Here, we introduce SynBioGPT v2.0 (https://synbiogpt.biodesign.ac.cn), which mitigates these limitations by decomposing queries into sub-questions and employing keyword-based searches. Tested on a 100-question synthetic biology benchmark, SynBioGPT v2.0 achieved 98% accuracy with the Claude-3.7-sonnet backend, a 10% improvement over v1.0s 88% with Llama3-8B-Instruct. This advance highlight the efficacy of query decomposition and precise retrieval in enhancing LLM utility for synthetic biology.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Datanator: an integrated database of molecular data for quantitatively modeling cellular behavior 92%
- PanKB: An interactive microbial pangenome knowledgebase for research, biotechnological innovation, and knowledge mining 92%
- CROssBAR: Comprehensive Resource of Biomedical Relations with Deep Learning Applications and Knowledge Graph Representations 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.