Assessing the ability of ChatGPT to extract natural product bioactivity and biosynthesis data from publications
Kalmer, T. L.; Ancajas, C. M. F.; Cheng, Z.; Oyedele, A. S.; Davis, H. L.; Walker, A.
10.1101/2024.08.01.606186 bioRxivShow abstract
Natural products are an excellent source of therapeutics and are often discovered through the process of genome mining, where genomes are analyzed by bioinformatic tools to determine if they have the biosynthetic capacity to produce novel or active compounds. Recently, several tools have been reported for predicting natural product bioactivities from the sequence of the biosynthetic gene clusters that produce them. These tools have the potential to accelerate the rate of natural product drug discovery by enabling the prioritization of novel biosynthetic gene clusters that are more likely to produce compounds with therapeutically relevant bioactivities. However, these tools are severely limited by a lack of training data, specifically data pairing biosynthetic gene clusters with activity labels for their products. There are many reports of natural product biosynthetic gene clusters and bioactivities in the literature that are not included in existing databases. Manual curation of these data is time consuming and inefficient. Recent developments in large language models and the chatbot interfaces built on top of them have enabled automatic data extraction from text, including scientific publications. We investigated how accurate ChatGPT is at extracting the necessary data for training models that predict natural product activity from biosynthetic gene clusters. We found that ChatGPT did well at determining if a paper described discovery of a natural product and extracting information about the products bioactivity. ChatGPT did not perform as well at extracting accession numbers for the biosynthetic gene cluster or producers genome although using an altered prompt improved accuracy.
Matching journals
The top 11 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- TargetDB: A target information aggregation tool and tractability predictor 92%
- Distinguishing classes of neuroactive drugs based on computational physicochemical properties and experimental phenotypic profiling in planarians 91%
- Exploring NCATS In-House Biomedical Data for Evidence-based Drug Repurposing 91%
Similar papers in this journal
- Atom Identifiers Generated by a Graph Coloring Method Enable Compound Harmonization Across Metabolic Databases 94%
- Benchmark dataset for training machine learning models to predict the pathway involvement of metabolites 93%
- Hierarchical Harmonization of Atom-Resolved Metabolic Re-actions Across Metabolic Databases 93%
Similar papers in this journal
Similar papers in this journal
- Predicting antimicrobial class specificity of small molecules using machine learning 94%
- The use of a graph database is a complementary approach to a classical similarity search for identifying commercially available fragment merges 92%
- Improving the reliability of molecular string representations for generative chemistry 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.