Investigations on using Evidence-Based GraphRag Pipeline using LLM Tailored for Answering USMLE Medical Exam Questions
Mohammed, S.; Fiaidhi, J.; Sekar, T.; Kushal, K.; Shankar, S.
Show abstract
The integration of evidence-based reasoning with retrieval-augmented generation (GraphRAG) holds great promise for enhancing large language model (LLM) question-answering (QA) capabilities. This research proposes a GraphRAG frame-work that improves the interpretability and reliability of LLM-generated answers in the medical domain. Our approach constructs a knowledge graph using Neo4j to represent UMLS medical entities and relationships, and complements it with a vector store of textbook embeddings for dense passage retrieval. The system is designed to combine symbolic reasoning and semantic search to produce more context-aware and evidence-grounded responses. As a proof of concept, we evaluate our system on United States Medical Licensing Examination (USMLE)-style questions, which require clinical reasoning across multiple domains. While overall answer accuracy remains comparable to that of an LLM-only baseline, our system consistently outperforms in citation fidelity -- providing richer, more traceable justifications by explicitly linking answers to graph paths and textbook passages. These findings suggest that even when correctness may vary, graph-informed retrieval improves transparency and auditability, which are critical for high-stakes domains like medicine. Our results motivate further refinement of hybrid GraphRAG systems to enhance both factual accuracy and clinical trustworthiness in QA applications.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Medication information extraction using local large language models 95%
- EHR-QC: A streamlined pipeline for automated electronic health records standardisation and preprocessing to predict clinical outcomes 95%
- De-novo FAIRification via an Electronic Data Capture system by automated transformation of filled electronic Case Report Forms into machine-readable data 95%
Similar papers in this journal
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 96%
- Ontology-based expansion of virtual gene panels to improve diagnostic efficiency for rare genetic diseases 94%
- MelAnalyze: Fact-Checking Melatonin claims using Large Language Models and Natural Language Inference 94%
Similar papers in this journal
Similar papers in this journal
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 95%
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 95%
- Comparative Effectiveness of Medical Concept Embedding for Feature Engineering in Phenotyping 95%
Similar papers in this journal
- One LLM is not Enough: Harnessing the Power of Ensemble Learning for Medical Question Answering 96%
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 95%
- Design and implementation of a system for automated monitoring of adherence to evidenced-based clinical guideline recommendations 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.