Back

Quantifying Evidence for Competing Biomedical Hypotheses using Large Language Models and Bayesian Analysis

Moore, B. M.; Freeman, J.; Millikin, R. J.; Mohanty, C.; George, K. S.; Bal, A.; Lock, C.; Sauer, J.-D.; Spurgeon, M. E.; Moore, D. L.; Travers, B. G.; Stewart, R.

2026-06-07 bioinformatics
10.64898/2026.06.05.730173 bioRxiv
Show abstract

Science fundamentally depends on the generation and testing of hypotheses, many of them controversial. An explosion in scientific literature has made evaluating hypotheses even within a domain a problem of scale, and risks slowing an already extensive consensus-building process. While this challenge has prompted interest in automated hypothesis evaluation tools, existing methods have not yet proven effective for comparing hypotheses. Here, we introduce KM-GPT-DCH, an algorithm that combines co-occurrence methods with large language models (LLMs) to develop a transparent and reproducible literature-based algorithm to compare controversial hypotheses using a structured scoring approach with Bayesian methods to estimate confidence. When testing the algorithm on historical controversial hypotheses previously decided, KM-GPT-DCH chooses the correct hypothesis with high confidence several years before the scientific community or public do so. We further apply the algorithm to compare twenty unresolved controversial hypothesis pairs providing guidance for future research. The method can help researchers and the public to evaluate biomedical hypotheses such as "Is it more likely that monoamine deficiency or inflammation causes depression?" It can also be used to assess and visualize historical trends in the scientific literature. A web-based implementation of the algorithm is freely available at https://skim.morgridge.org.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Bioinformatics Advances
203 papers in training set
Top 0.1%
18.4%
2
Bioinformatics
1204 papers in training set
Top 2%
15.0%
3
BMC Bioinformatics
457 papers in training set
Top 1.0%
7.9%
4
Genome Biology
637 papers in training set
Top 2%
5.5%
5
BioData Mining
22 papers in training set
Top 0.1%
5.5%
50% of probability mass above
6
PLOS Computational Biology
1863 papers in training set
Top 8%
4.3%
7
NAR Genomics and Bioinformatics
242 papers in training set
Top 1.0%
4.3%
8
Cell Reports Methods
165 papers in training set
Top 0.7%
3.2%
9
Briefings in Bioinformatics
354 papers in training set
Top 3%
3.2%
10
PLOS ONE
5266 papers in training set
Top 38%
3.2%
11
Computational and Structural Biotechnology Journal
242 papers in training set
Top 2%
2.4%
12
Database
61 papers in training set
Top 0.3%
2.4%
13
BMC Genomics
406 papers in training set
Top 4%
1.7%
14
Nucleic Acids Research
1281 papers in training set
Top 9%
1.7%
15
GigaScience
212 papers in training set
Top 3%
1.7%
16
Scientific Reports
3612 papers in training set
Top 60%
1.4%
17
eLife
5828 papers in training set
Top 57%
1.1%
18
Nature Communications
5641 papers in training set
Top 51%
1.1%
19
PeerJ
308 papers in training set
Top 9%
1.0%
20
PLOS Biology
486 papers in training set
Top 10%
1.0%
21
Patterns
78 papers in training set
Top 3%
0.8%
22
The American Journal of Human Genetics
234 papers in training set
Top 3%
0.8%
23
iScience
1154 papers in training set
Top 40%
0.6%