Back

Multi-Agent Dynamic Refinement Outperforms Static RAG in Clinical Reasoning for Complex Nephrology Cases

Yano, Y.; Kakizaki, H.; Nagasu, H.; Kishi, S.; Koshida, T.; Nihei, Y.; Hirano, A.; Sugawara, Y.; Imaizumi, T.; Osakabe, Y.; Sakaguchi, Y.; Nangaku, M.; Mori, H.; Naito, T.; Ohashi, M.; Maruyama, S.; Matsui, I.; Isaka, Y.; Okada, H.; Suzuki, Y.; Kashihara, N.

2026-07-16 nephrology
10.64898/2026.07.15.26358121 medRxiv
Show abstract

Background: Large language models (LLMs) struggle with dynamic, longitudinal clinical reasoning. We developed a Multi-Stage Iterative Clinical Reasoning Agent framework to address this gap and systematically decouple the clinical efficacy of static retrieval-augmented generation (RAG) from dynamic self-refinement. Methods: Ten complex longitudinal nephrology cases, rigorously selected via a modified Delphi consensus technique, were blindly evaluated by four board-certified nephrologists and a multi-model AI panel. We compared three architectures across nine cognitive steps: (Model A) a baseline frontier LLM, (Model B) an LLM augmented with static guideline-based RAG, and (Model C) our proposed multi-agent framework featuring RAG integrated with iterative self-critique and refinement. Results: In human evaluations (20-point scale), Model C (mean 17.2, SD 1.2) significantly outperformed both Model A (16.1, 1.3) and Model B (16.2, 1.2) (P < 0.001). Implementing static RAG (Model B) yielded no significant improvement over the baseline. Automated AI evaluations (15-point scale) corroborated these findings: Model C (14.7, 0.6) outscored Model A (14.2, 0.9, P < 0.001) and Model B (14.3, 0.9, P = 0.01). While monolithic models exhibited severe score degradations in planning-heavy tasks such as dynamic differential diagnoses, the multi-agent framework effectively intercepted error cascades, achieving significantly higher diagnostic accuracy (mean 17.6, P = 0.019) and therapeutic management scores (17.3, P = 0.002). Conclusions: Static knowledge retrieval alone fails to enhance frontier LLM performance in longitudinal medical reasoning. Distributing clinical workflows into a multi-agent dynamic refinement pipeline significantly improves reasoning completeness, intercepts error cascades, and safely resolves planning bottlenecks in complex patient care.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

1
PLOS ONE
5266 papers in training set
Top 14%
13.1%
2
npj Digital Medicine
118 papers in training set
Top 0.7%
8.0%
3
Communications Medicine
113 papers in training set
Top 0.2%
8.0%
4
Journal of the American Medical Informatics Association
71 papers in training set
Top 0.5%
6.4%
5
Scientific Reports
3612 papers in training set
Top 22%
4.4%
6
iScience
1154 papers in training set
Top 4%
4.1%
7
JAMA Network Open
130 papers in training set
Top 0.8%
4.1%
8
Clinical Journal of the American Society of Nephrology
10 papers in training set
Top 0.1%
4.1%
50% of probability mass above
9
Kidney360
22 papers in training set
Top 0.2%
3.3%
10
International Journal of Medical Informatics
26 papers in training set
Top 0.4%
2.8%
11
BMC Medicine
176 papers in training set
Top 1%
2.8%
12
Frontiers in Digital Health
24 papers in training set
Top 0.5%
2.5%
13
BMC Nephrology
18 papers in training set
Top 0.2%
2.4%
14
Journal of Clinical Medicine
97 papers in training set
Top 2%
2.2%
15
PLOS Digital Health
106 papers in training set
Top 2%
2.2%
16
Frontiers in Medicine
120 papers in training set
Top 2%
1.9%
17
PLOS Global Public Health
344 papers in training set
Top 5%
1.9%
18
BMC Medical Informatics and Decision Making
43 papers in training set
Top 1.0%
1.8%
19
Journal of Biomedical Informatics
47 papers in training set
Top 0.8%
1.5%
20
Kidney International Reports
15 papers in training set
Top 0.2%
1.4%
21
Life
29 papers in training set
Top 0.4%
1.1%
22
Bioinformatics Advances
203 papers in training set
Top 4%
1.1%
23
JAMIA Open
42 papers in training set
Top 1%
1.1%
24
JMIR Medical Informatics
18 papers in training set
Top 0.7%
1.1%
25
Journal of Translational Medicine
57 papers in training set
Top 2%
1.0%
26
Medicine
31 papers in training set
Top 2%
0.9%
27
Frontiers in Public Health
148 papers in training set
Top 6%
0.9%
28
Healthcare
17 papers in training set
Top 1.0%
0.9%
29
Health Science Reports
13 papers in training set
Top 0.7%
0.6%
30
Canadian Medical Association Journal
15 papers in training set
Top 0.2%
0.6%