Back

Creating a biomedical knowledge base by addressing GPT inaccurate responses and benchmarking context

Darnell, S. S.; Prins, J. P.; Suh, E.; Huang, P.; Williams, R. W.; Overall, R.; Chen, H.; Garrison, E.; Guarracino, A.; Villani, F.; Muli, P.; Ashbrook, D. G.; Colonna, V.; Batten, C.; Sen, S.; Muriithi, F. M.; Yousefi, S.; Nijveen, H.; Lisso, F.; Isaac, A.; Kabui, A.; Kilungi, M. B.; Kibet, A.; Umar, M.; Muhia, B.

2024-10-18 scientific communication and education
10.1101/2024.10.16.618663 bioRxiv
Show abstract

We created GNQA, a generative pre-trained transformer (GPT) knowledge base driven by a performant retrieval augmented generation (RAG) with a focus on aging, dementia, Alzheimers and diabetes. We uploaded a corpus of three thousand peer reviewed publications on these topics into the RAG. To address concerns about inaccurate responses and GPT hallucinations, we implemented a context provenance tracking mechanism that enables researchers to validate responses against the original material and to get references to the original papers. To assess the effectiveness of contextual information we collected evaluations and feedback from both domain expert users and citizen scientists on the relevance of GPT responses. A key innovation of our study is automated evaluation by way of a RAG assessment system (RAGAS). RAGAS combines human expert assessment with AI-driven evaluation to measure the effectiveness of RAG systems. When evaluating the responses to their questions, human respondents give a "thumbs-up" 76% of the time. Meanwhile, RAGAS scores 90% on answer relevance on questions posed by experts. And when GPT-generates questions, RAGAS scores 74% on answer relevance. With RAGAS we created a benchmark that can be used to continuously assess the performance of our knowledge base. Full GNQA functionality is embedded in the free GeneNetwork.org web service, an open-source system containing over 25 years of experimental data on model organisms and human. The code developed for this study is published under a free and open-source software license at https://git.genenetwork.org/gn-ai/tree/README.md.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Bioinformatics Advances
203 papers in training set
Top 0.1%
27.2%
2
PLOS ONE
5266 papers in training set
Top 17%
11.3%
3
GigaScience
212 papers in training set
Top 0.1%
10.0%
4
PLOS Computational Biology
1863 papers in training set
Top 5%
6.9%
50% of probability mass above
5
Bioinformatics
1204 papers in training set
Top 4%
6.4%
6
Heliyon
152 papers in training set
Top 2%
2.7%
7
Patterns
78 papers in training set
Top 0.8%
2.4%
8
Frontiers in Artificial Intelligence
20 papers in training set
Top 0.2%
2.4%
9
Scientific Reports
3612 papers in training set
Top 46%
2.2%
10
BioData Mining
22 papers in training set
Top 0.2%
1.8%
11
BMC Bioinformatics
457 papers in training set
Top 4%
1.8%
12
Frontiers in Bioinformatics
49 papers in training set
Top 0.5%
1.5%
13
eneuro
439 papers in training set
Top 5%
1.4%
14
iScience
1154 papers in training set
Top 24%
1.2%
15
eLife
5828 papers in training set
Top 56%
1.2%
16
JAMIA Open
42 papers in training set
Top 1%
1.1%
17
F1000Research
88 papers in training set
Top 3%
1.0%
18
Ecological Informatics
33 papers in training set
Top 0.7%
0.9%
19
Journal of Biomedical Informatics
47 papers in training set
Top 1%
0.9%
20
Gigabyte
62 papers in training set
Top 1%
0.9%
21
Journal of Cheminformatics
29 papers in training set
Top 0.7%
0.9%
22
Ecology and Evolution
267 papers in training set
Top 5%
0.9%
23
Database
61 papers in training set
Top 1%
0.6%
24
Frontiers in Physiology
106 papers in training set
Top 3%
0.6%
25
Frontiers in Systems Biology
10 papers in training set
Top 0.2%
0.6%
26
Lab Animal
11 papers in training set
Top 0.3%
0.6%
27
Journal of the American Medical Informatics Association
71 papers in training set
Top 2%
0.6%