RAGnosis: Retrieval-Augmented Generation for Enhanced Medical Decision Making
Rouhollahi, A.; Homaei, A.; Sahu, A.; Harari, R.; Nezami, F. R.
Show abstract
We present RAGnosis, a fully offline, retrieval-augmented framework for interpreting unstructured clinical text using open-weight large language models. As a proof of concept, we apply RAGnosis to the task of paravalvular leak (PVL) classification from cardiac catheterization reports, a process that typically requires slow, expert-driven interpretation. The system combines local OCR, semantic retrieval, and instruction-tuned LLMs to generate evidence-backed predictions and explanations grounded in real clinical documentation. We evaluate four models (DeepSeek 70B, Gemma 27B, Mistral 7B, LLaMA 3B) across 100 reports, analyzing classification accuracy, explanation quality, and retrieval relevance. Results highlight tradeoffs between fluency and reliability, with DeepSeek demonstrating the most consistent performance. By operating entirely on-prem and supporting modular integration, RAGnosis provides a scalable and interpretable foundation for clinical NLP that delivers not just answers but traceable reasoning.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 96%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 95%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 95%
Similar papers in this journal
Similar papers in this journal
- Evaluating Anti-LGBTQIA+ Medical Bias in Large Language Models 94%
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 94%
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 93%
Similar papers in this journal
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 95%
- Design and Implementation of an End-to-End AI-Driven Colonoscopy Recall Workflow at Scale 94%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.