Back

Optimizing the Clinical Application of Rheumatology Guidelines Using Large Language Models: A Retrieval-Augmented Generation Framework Integrating ACR and EULAR Recommendations

Garcia, A. M.; Benavent, D.; Barbancho, B. M.; Nunez, D. F.

2025-04-11 rheumatology
10.1101/2025.04.10.25325588 medRxiv
Show abstract

ObjectivesTimely access to current rheumatology guidelines at the point of care is challenging. We aimed to develop and evaluate the first Retrieval-Augmented Generation (RAG) system specifically designed for adult rheumatology, integrating European Alliance of Associations for Rheumatology (EULAR) and American College of Rheumatology (ACR) guidelines to provide rheumatologists with timely, evidence-based recommendations at the point of care. MethodsEULAR and ACR management guidelines were selected by rheumatologists based on their clinical relevance for decision making and processed. A RAG system was implemented using LangChain framework, voyage-3 embedding model, and a Qdrant vector database. To evaluate the system, ten questions per guideline were generated using ChatGPT 4.5. Answers to these guideline-specific questions were subsequently produced by ChatGPT-o3-mini with context retrieval (RAG) and without (baseline). Performance was assessed by an LLM-as-a-judge (Gemini 2.0 Flash) using a 5-point Likert scale across five dimensions: relevance, factual accuracy, safety, completeness, and conciseness. The judge also determined preference between the RAG and baseline responses. Statistical significance was established using Wilcoxon signed-rank and Binomial tests. For further validation, two blinded rheumatologists independently evaluated a random sample of questions (15%). ResultsAfter agreement, 74 guidelines were included, and 740 evaluation questions were generated. Analysis revealed that the RAG system significantly outperformed the baseline system across all criteria (p<0.001) in the LLM-as-a-judge evaluation. Manual evaluation by rheumatologists confirmed these findings (p<0.001 for accuracy, safety, completeness). Furthermore, the RAG system was significantly preferred by the LLM-as-a-judge in 92.8% of comparisons (p<0.001) and by the human evaluators in 71.2%-74.8% of comparisons (p<0.001). ConclusionThis study demonstrates the successful development and evaluation of a RAG system integrating extensive EULAR/ACR guidelines for adult rheumatology. The system significantly improves answer quality compared to a baseline LLM. This provides a robust foundation for reliable, AI-driven clinical decision support tools designed to enhance guideline adherence and evidence-based practice in rheumatology by providing clinicians with rapid, context-aware access to recommendations. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=104 SRC="FIGDIR/small/25325588v2_ufig1.gif" ALT="Figure 1"> View larger version (43K): org.highwire.dtl.DTLVardef@1757c6forg.highwire.dtl.DTLVardef@3c78e1org.highwire.dtl.DTLVardef@2424eeorg.highwire.dtl.DTLVardef@f49ef9_HPS_FORMAT_FIGEXP M_FIG C_FIG Key messagesO_LILarge language models, combined with EULAR and ACR guidelines, may enhance rheumatology clinical decision support. C_LIO_LIRetrieval augmented generation (RAG) responses showed significantly greater accuracy, safety and completeness than baseline LLMs. C_LIO_LIRAG is a promising architecture for reducing hallucinations and providing grounded, reliable answers. C_LI

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.