Enzyme Co-Scientist: Harnessing Large Language Models for Enzyme Kinetic Data Extraction from Literature
Jiang, J.; Hu, J.; Xie, S.; Guo, M.; Dong, Y.; Fu, S.; Jiang, X.; Yue, Z.; Shi, J.; Zhang, X.; Song, M.; Chen, G.; Lu, H.; Wu, X.; Guo, P.; Han, D.; Sun, Z.; Qiu, J.
Show abstract
The extraction of molecular annotations from scientific literature is critical for advancing data-driven research. However, traditional methods, which primarily rely on human curation, are labor-intensive and error-prone. Here, we present an LLM-based agentic workflow that enables automatic and efficient data extraction from literature with high accuracy. As a demonstration, our workflow successfully delivers a dataset containing over 91,000 enzyme kinetics entries from around 3,500 papers. It achieves an average F1 score above 0.9 on expert-annotated subsets of protein enzymes and can be extended to the ribozyme domain in fewer than 3 days at less than $90. This method opens up new avenues for accelerating the pace of scientific research.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CROssBAR: Comprehensive Resource of Biomedical Relations with Deep Learning Applications and Knowledge Graph Representations 95%
- Datanator: an integrated database of molecular data for quantitatively modeling cellular behavior 93%
- Thunor: Visualization and Analysis of High-Throughput Dose-response Datasets 92%
Similar papers in this journal
- AutoPM3: Enhancing Variant Interpretation via LLM-driven PM3 Evidence Extraction from Scientific Literature 95%
- GOFlowLLM - Curating miRNA literature with Large Language Models and flowcharts 95%
- FuncFetch: An LLM-assisted workflow enables mining thousands of enzyme-substrate interactions from published manuscripts 95%
Similar papers in this journal
- biocentral: embedding-based protein predictions 93%
- RNA3DB: a structurally-dissimilar dataset split for training and benchmarking deep learning models for RNA structure prediction 92%
- Effectiveness and efficiency: label-aware hierarchical subgraph learning for protein-protein interaction (laruGL-PPI) 92%
Similar papers in this journal
Similar papers in this journal
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database 94%
- PEPhub: a database, web interface, and API for editing, sharing, and validating biological sample metadata 93%
- TooManyCellsInteractive: a visualization tool for dynamic exploration of single-cell data 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.