Back

Evaluation of search-enabled Pre-trained Large Language Models on retrieval tasks for the PubChem Database

Sze, A.; Hassoun, S.

2024-08-19 bioinformatics
10.1101/2024.08.15.608120 bioRxiv
Show abstract

Databases are indispensable in biological and biomedical research, hosting vast amounts of structured and unstructured data, facilitating the organization, retrieval, and analysis of complex data. Database access, however, remains a manual, tedious, and sometimes overwhelming, task. We investigate in this study the current state of using pre-trained, search-enabled LLMs for data retrieval from biological databases. Equipped with internet search and code generation capabilities, LLMs promise to streamline database access through natural language, expedite search and knowledge retrieval, and provide coherent analytical summaries. As an example database, we focus on evaluating a current search-enabled LLMs (GPT-4o) for retrieval from the PubChem database, a flagship, heavily used database that plays a critical role in biological and biomedical research. As PubChem is an open archival repository, it provides a well-documented programmatic interface that can be exploited through LLM code generation capabilities. We evaluate retrieval tasks for eight common PubChem access protocols that were previously documented. The tasks include identifying interacting genes and proteins, finding drug-like compounds based on structural similarity, retrieving bioactivity data, and locating stereoisomers and isotopomers. We develop a methodology for adopting the protocols into an LLM-prompt, where we supplement the prompt with additional context through iterative prompt refinement as needed. To further evaluate the LLM capabilities, we instruct the LLM to perform the retrieval with and without using programmatic access. We compare the results (referred to as gold and silver answers) when using these retrieval modalities with two traditional retrieval baselines that include running the manual search steps for each reference protocol through the PubChem database web interface, and through the provided PUG (Power-User Gateway) programmatic access. We quantitatively and qualitatively summarize our results, showing that generating programmatic access is more likely to yield the correct answers. We highlight the value and limitations of using current search-based LLMs for database retrieval. We also provide guidance for the future development that can improve the accuracy and reliability of search-based LLMs.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Journal of Cheminformatics
29 papers in training set
Top 0.1%
22.5%
2
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.3%
18.6%
3
PLOS ONE
5266 papers in training set
Top 24%
6.8%
4
Bioinformatics
1204 papers in training set
Top 3%
6.8%
50% of probability mass above
5
Bioinformatics Advances
203 papers in training set
Top 1%
4.1%
6
Computational and Structural Biotechnology Journal
242 papers in training set
Top 1%
3.2%
7
Database
61 papers in training set
Top 0.2%
3.2%
8
Metabolites
53 papers in training set
Top 0.3%
3.2%
9
BMC Bioinformatics
457 papers in training set
Top 3%
2.4%
10
PLOS Computational Biology
1863 papers in training set
Top 13%
1.9%
11
Artificial Intelligence in the Life Sciences
13 papers in training set
Top 0.1%
1.7%
12
Molecules
39 papers in training set
Top 0.7%
1.5%
13
Nucleic Acids Research
1281 papers in training set
Top 10%
1.3%
14
Briefings in Bioinformatics
354 papers in training set
Top 6%
1.1%
15
Scientific Reports
3612 papers in training set
Top 68%
1.1%
16
Journal of Molecular Biology
232 papers in training set
Top 3%
1.1%
17
Frontiers in Pharmacology
111 papers in training set
Top 2%
1.0%
18
Nature Protocols
33 papers in training set
Top 0.4%
1.0%
19
International Journal of Molecular Sciences
494 papers in training set
Top 13%
1.0%
20
BioData Mining
22 papers in training set
Top 0.6%
1.0%
21
Frontiers in Bioinformatics
49 papers in training set
Top 1%
0.8%
22
Journal of the American Medical Informatics Association
71 papers in training set
Top 2%
0.6%
23
Biomolecules
100 papers in training set
Top 3%
0.6%
24
JAMIA Open
42 papers in training set
Top 2%
0.6%
25
eLife
5828 papers in training set
Top 68%
0.6%