Development and validation of Retrieval Augmented Generation (RAG) and GraphRAG for complex clinical cases
Sesen, B.; Au Yeung, J.; Asgari, E.
Show abstract
ObjectiveChronic Kidney Disease (CKD) is a progressive condition requiring evidence-based management, but adherence to complex guidelines remains challenging. Large Language Models (LLMs) could support clinical decision-making, yet their unreliability limits direct use. This study aimed to evaluate whether Retrieval-Augmented Generation (RAG), particularly a knowledge graph-enhanced pipeline (GraphRAG), improves guideline-based clinical decision support (CDS) in CKD management. Methods and AnalysisWe compared three approaches: a baseline LLM (GPT-4o), a vector-indexed RAG pipeline, and a GraphRAG pipeline. Each model answered nine clinically relevant questions for a synthetic cohort of 70 CKD patients. Outputs were assessed for clinical correctness, patient-specificity, and clarity, using both clinician-led evaluations and an LLM-as-Judge framework. ResultsRAG-based methods outperformed the baseline LLM in clinical correctness and guideline adherence. GraphRAG achieved the highest patient-specificity by leveraging multi-hop relationships across a knowledge graph derived from NICE CKD guidelines, particularly for tasks involving thresholds, algorithmic decisions, or open-ended management. However, GraphRAG scored lower in clarity, as its graph walks often returned long guideline excerpts that obscured key recommendations. All RAG systems were limited by the scope of the indexed guideline and performed poorly when essential information was missing. ConclusionsRAG and GraphRAG provide a scalable, auditable foundation for guideline-aligned CDS in CKD, with GraphRAG showing particular strengths in tailoring advice to patient data. Nonetheless, trade-offs remain between specificity and clarity, and effective deployment will require robust content management, transparent validation pipelines, and integration within established clinical governance frameworks. Key points- LLMs have comprehensive medical knowledge but require access to up-to-date, evidence-based, and locally relevant guidelines to be effective in CDS. - Hallucinations (the generation of inaccurate or misleading information) remain a major limitation for LLMs in healthcare. - Traditional information retrieval methods face several challenges in providing accurate, context-specific evidence. - Retrieval-Augmented Generation (RAG) and graph-based RAG approaches have emerged as promising solutions to overcome these limitations. - Renal medicine provides an ideal test domain to evaluate these models, given its complexity and reliance on nuanced, multidisciplinary decision-making. - Studying LLM performance in kidney health can yield valuable insights into how such models can safely and effectively support complex clinical decision-making.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 95%
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 94%
- Causal modeling of chronic kidney disease in a participatory framework for informing the inclusion of social drivers in health algorithms 93%
Similar papers in this journal
- Extraction of Crohn's Disease Clinical Phenotypes from Clinical Text Using Natural Language Processing 93%
- Achieving Inclusive Healthcare through Integrating Education and Research with AI and Personalized Curricula 93%
- Systematic Review of Large Language Models for Patient Care: Current Applications and Challenges 92%
Similar papers in this journal
- Clinical code sets and the problem of redundancy in code set repositories 92%
- A machine learning-based phenotype for long COVID in children: an EHR-based study from the RECOVER program 92%
- ChatGPT- versus human-generated answers to frequently asked questions about diabetes: a Turing test-inspired survey among employees of a Danish diabetes center 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.