Back

Enzyme Co-Scientist: Harnessing Large Language Models for Enzyme Kinetic Data Extraction from Literature

Jiang, J.; Hu, J.; Xie, S.; Guo, M.; Dong, Y.; Fu, S.; Jiang, X.; Yue, Z.; Shi, J.; Zhang, X.; Song, M.; Chen, G.; Lu, H.; Wu, X.; Guo, P.; Han, D.; Sun, Z.; Qiu, J.

2025-03-11 biochemistry
10.1101/2025.03.03.641178 bioRxiv
Show abstract

The extraction of molecular annotations from scientific literature is critical for advancing data-driven research. However, traditional methods, which primarily rely on human curation, are labor-intensive and error-prone. Here, we present an LLM-based agentic workflow that enables automatic and efficient data extraction from literature with high accuracy. As a demonstration, our workflow successfully delivers a dataset containing over 91,000 enzyme kinetics entries from around 3,500 papers. It achieves an average F1 score above 0.9 on expert-annotated subsets of protein enzymes and can be extended to the ribozyme domain in fewer than 3 days at less than $90. This method opens up new avenues for accelerating the pace of scientific research.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.