A Clinically-Informed Framework for Evaluating Vision-Language Models in Radiology Report Generation: Taxonomy of Errors and Risk-Aware Metric
Guan, H.; Hou, P. C.; Hong, P.; Wang, L.; Zhang, W.; Du, X.; Zhou, Z.; Zhou, L.
Show abstract
Recent advances in vision-language models (VLMs) have enabled automatic radiology report generation, yet current evaluation methods remain limited to general-purpose NLP metrics or coarse classification-based clinical scores. In this study, we propose a clinically informed evaluation framework for VLM-generated radiology reports that goes beyond traditional performance measures. We define a taxonomy of 12 radiology-specific error types, each annotated with clinical risk levels (low, medium, high) in collaboration with physicians. Using this framework, we conduct a comprehensive error analysis of three representative VLMs, i.e., DeepSeek VL2, CXR-LLaVA, and CheXagent, on 685 gold-standard, expert-annotated MIMIC-CXR cases. We further introduce a risk-aware evaluation metric, the Clinical Risk-weighted Error Score for Text-generation (CREST), to quantify safety impact. Our findings reveal critical model vulnerabilities, common error patterns, and condition-specific risk profiles, offering actionable insights for model development and deployment. This work establishes a safety-centric foundation for evaluating and improving medical report generation models. The source code of our evaluation framework, including CREST computation and error taxonomy analysis, is available at https://github.com/guanharry/VLM-CREST.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- EchoGraph System for Automated Quality Assessment of Echocardiography Reports 94%
- The clinician-AI interface: intended use and explainability in FDA-cleared AI devices for medical image interpretation 94%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
Similar papers in this journal
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 95%
- Development and Validation of a Deep Learning Model for Detecting Signs of Tuberculosis on Chest Radiographs among US-bound Immigrants and Refugees 94%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 94%
Similar papers in this journal
- ARDSFlag: An NLP/Machine Learning Algorithm to Visualize and Detect High-Probability ARDS Admissions Independent of Provider Recognition and Billing Codes 94%
- Equipping Computational Pathology Systems with Artifact Processing Pipelines: A Showcase for Computation and Performance Trade-offs 93%
- Addressing Label Noise for Electronic Health Records: Insights from Computer Vision for Tabular Data 92%
Similar papers in this journal
- Toward Understanding COVID-19 Pneumonia: A Deep-learning-based Approach for Severity Analysis and Monitoring the Disease 94%
- Tracking And Predicting COVID-19 Radiological Trajectory Using Deep Learning On Chest X-Rays: Initial Accuracy Testing 94%
- On evaluation metrics for medical applications of artificial intelligence 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.