Enhancing Clinical Reasoning In Medical Vision-Language Model Through Structured Prompts
Dasaramoole Prakash, K.; Han, Y.; Kim, K.
Show abstract
Medical Vision-Language Models (MVLMs) are emerging as powerful tools for tasks such as Visual Question Answering (VQA); however, they often struggle with hallucination and limited reasoning transparency, particularly in complex diagnostic scenarios. In this work, we enhance the MedVLM-R1 framework by fine-tuning it using clinically informed prompt structures tailored specifically for radiology-based reasoning. Without altering the original model architecture or training strategy, we redesign the system prompts and question templates to guide the model through structured, modality-aware, and step-by-step diagnostic reasoning. Fine-tuning is performed using MRI-based question-answer (QA) pairs, and evaluations are conducted across three diagnostic imaging: MRI, CT, and X-ray to assess both in-domain and out-of-domain generalization. Our approach improves reasoning transparency and accuracy, achieving 96.00% on MRI, 72.67% on CT, and 75.2% on X-ray. Compared to the original MedVLM-R1, our method closes the gap in MRI accuracy while significantly enhancing generalization performance on CT and X-ray modalities. These results demonstrate that clinically grounded prompting effectively improves both reasoning fidelity and robustness across imaging modalities. The code is available at our GitHub repository:https://github.com/aidanbio/AIdanMed
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 94%
- A human-in-the-loop explanation framework for morphologically transparent AI predictions from whole-slide images 93%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 92%
Similar papers in this journal
- LCD Benchmark: Long Clinical Document Benchmark on Mortality Prediction for Language Models 93%
- ENRICHing Medical Imaging Training Sets Enables More Efficient Machine Learning 92%
- Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning models 92%
Similar papers in this journal
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 94%
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 94%
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 92%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.