Clinical selectivity and failure modes of automated chest radiograph report evaluation metrics: a cross-dataset analysis of ReXErr-v1 and RadEvalX
Naidu, J.; Muralidharan, S.; Prashani, A.; Baskaradoss, V.
Show abstract
Objectives: To test whether radiology report evaluation metrics distinguish clinically meaningful errors from textual changes and align with radiologist-assessed error burden. Methods: Cross-dataset evaluation used ReXErr-v1 (2,708 report pairs; 5,724 paired error sentences) and 100 RadEvalX report pairs with expert error counts. BLEU-4, ROUGE-L and METEOR were assessed in ReXErr-v1; RadEvalX analyses included these plus BERTScore, CheXbert, RadGraph F1 and RadCliQ. Outcomes were ReXErr-v1 pairwise win rate and AUROC for clinical-content versus linguistic errors, and RadEvalX Spearman correlation with clinically significant error count and AUROC for any significant error. Confidence intervals used 10,000 clustered percentile bootstrap resamples; Holm adjustment-controlled multiplicity. Results: ReXErr-v1 paired-sentence win rates were 0.986 for BLEU-4, 0.999 for ROUGE-L and 0.998 for METEOR, but discrimination of clinical-content from linguistic errors was modest (AUROC 0.609-0.620). Penalty magnitude was strongly associated with textual change after adjustment for error type (normalised character edit distance coefficient 0.746; 95% CI 0.705-0.788; P<0.001). In RadEvalX, CheXbert showed the highest correlation with clinically significant errors (rho=0.413; 95% CI 0.223-0.578) and highest AUROC (0.742; 95% CI 0.638-0.836). Conclusions: Near-ceiling sensitivity to textual corruption did not imply sensitivity to clinical significance. CheXbert showed the highest alignment with expert error assessment, although pairwise superiority was not demonstrated over all comparators and performance remained moderate.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Observer agreement and clinical significance of chest CT reporting in patients suspected of COVID-19 93%
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 93%
- Evaluating Large Language Model-Generated Brain MRI Protocols: Performance of GPT4o, o3-mini, DeepSeek-R1 and Qwen2.5-72B 91%
Similar papers in this journal
- tbiExtractor: A framework for Extracting Traumatic Brain Injury Common Data Elements from Radiology Reports 91%
- A method for rapid machine learning development for data mining with Doctor-In-The-Loop 90%
- Enhancing Semantic Segmentation in Chest X-Ray Images through Image Preprocessing: ps-KDE for Pixel-wise Substitution by Kernel Density Estimation 90%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Generative Large Language Models in Electronic Health Records for Patient Care Since 2023: A Systematic Review 90%
- Use of unstructured text in prognostic clinical prediction models: a systematic review 90%
- Is One Run Enough? Reproducibility of Flagship Large Language Models Across Temperature and Reasoning Settings in Biomedical Text Processing 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.