Optimizing the Clinical Application of Rheumatology Guidelines Using Large Language Models: A Retrieval-Augmented Generation Framework Integrating ACR and EULAR Recommendations
Garcia, A. M.; Benavent, D.; Barbancho, B. M.; Nunez, D. F.
Show abstract
ObjectivesTimely access to current rheumatology guidelines at the point of care is challenging. We aimed to develop and evaluate the first Retrieval-Augmented Generation (RAG) system specifically designed for adult rheumatology, integrating European Alliance of Associations for Rheumatology (EULAR) and American College of Rheumatology (ACR) guidelines to provide rheumatologists with timely, evidence-based recommendations at the point of care. MethodsEULAR and ACR management guidelines were selected by rheumatologists based on their clinical relevance for decision making and processed. A RAG system was implemented using LangChain framework, voyage-3 embedding model, and a Qdrant vector database. To evaluate the system, ten questions per guideline were generated using ChatGPT 4.5. Answers to these guideline-specific questions were subsequently produced by ChatGPT-o3-mini with context retrieval (RAG) and without (baseline). Performance was assessed by an LLM-as-a-judge (Gemini 2.0 Flash) using a 5-point Likert scale across five dimensions: relevance, factual accuracy, safety, completeness, and conciseness. The judge also determined preference between the RAG and baseline responses. Statistical significance was established using Wilcoxon signed-rank and Binomial tests. For further validation, two blinded rheumatologists independently evaluated a random sample of questions (15%). ResultsAfter agreement, 74 guidelines were included, and 740 evaluation questions were generated. Analysis revealed that the RAG system significantly outperformed the baseline system across all criteria (p<0.001) in the LLM-as-a-judge evaluation. Manual evaluation by rheumatologists confirmed these findings (p<0.001 for accuracy, safety, completeness). Furthermore, the RAG system was significantly preferred by the LLM-as-a-judge in 92.8% of comparisons (p<0.001) and by the human evaluators in 71.2%-74.8% of comparisons (p<0.001). ConclusionThis study demonstrates the successful development and evaluation of a RAG system integrating extensive EULAR/ACR guidelines for adult rheumatology. The system significantly improves answer quality compared to a baseline LLM. This provides a robust foundation for reliable, AI-driven clinical decision support tools designed to enhance guideline adherence and evidence-based practice in rheumatology by providing clinicians with rapid, context-aware access to recommendations. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=104 SRC="FIGDIR/small/25325588v2_ufig1.gif" ALT="Figure 1"> View larger version (43K): org.highwire.dtl.DTLVardef@1757c6forg.highwire.dtl.DTLVardef@3c78e1org.highwire.dtl.DTLVardef@2424eeorg.highwire.dtl.DTLVardef@f49ef9_HPS_FORMAT_FIGEXP M_FIG C_FIG Key messagesO_LILarge language models, combined with EULAR and ACR guidelines, may enhance rheumatology clinical decision support. C_LIO_LIRetrieval augmented generation (RAG) responses showed significantly greater accuracy, safety and completeness than baseline LLMs. C_LIO_LIRAG is a promising architecture for reducing hallucinations and providing grounded, reliable answers. C_LI
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 94%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 94%
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 93%
Similar papers in this journal
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 93%
- Imbalanced Machine Learning Classification Models For Removal Biosimilar Drugs And Increased Activity In Patients With Rheumatic Diseases 92%
- An interactive retrieval system for clinical trial studies with context-dependent protocol elements 92%
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 93%
- The clinician-AI interface: intended use and explainability in FDA-cleared AI devices for medical image interpretation 92%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 92%
Similar papers in this journal
- Computer vision detects inflammatory arthritis in standardized smartphone photographs in an Indian patient cohort 93%
- Streamlining Intersectoral Provision of Real-World Health Data: A Service Platform for Improved Clinical Research and Patient Care 92%
- Emerging Applications of NLP and Large Language Models in Gastroenterology and Hepatology: A Systematic Review 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.