RaTEScore: A Metric for Radiology Report Generation
Zhao, W.; Wu, C.; Zhang, X.; Zhang, Y.; Wang, Y.; Xie, W.
Show abstract
This paper introduces a novel, entity-aware metric, termed as Radiological Report (Text) Evaluation (RaTEScore), to assess the quality of medical reports generated by AI models. RaTEScore emphasizes crucial medical entities, such as diagnostic outcomes and anatomical details. Moreover, it is robust against medical synonyms and sensitive to negation expressions. Technically, we developed a comprehensive medical NER dataset, RaTE-NER, and trained an NER model specifically for this purpose. This model enables the decomposition of complex radiological reports into constituent medical entities. The metric itself is derived by comparing the similarity of entity embeddings, obtained from a language model, based on their types and relevance to clinical significance. Our evaluations demonstrate that RaTEScore aligns more closely with human preference than existing metrics, validated both on established public benchmarks and our newly proposed RaTE-Eval benchmark.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Effective Deep Learning Approaches for Predicting COVID-19 Outcomes from Chest Computed Tomography Volumes 94%
- Assisting Scalable Diagnosis Automatically via CT Images in the Combat against COVID-19 93%
- Toward Understanding COVID-19 Pneumonia: A Deep-learning-based Approach for Severity Analysis and Monitoring the Disease 92%
Similar papers in this journal
- EchoGraph System for Automated Quality Assessment of Echocardiography Reports 92%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 91%
- A human-in-the-loop explanation framework for morphologically transparent AI predictions from whole-slide images 91%
Similar papers in this journal
- Building a Best-in-Class De-identification Tool for Electronic Medical Records Through Ensemble Learning 94%
- Structuring clinical text with AI: old vs. new natural language processing techniques evaluated on eight common cardiovascular diseases 91%
- Inferring global-scale temporal latent topics from news reports to predict public health interventions for COVID-19 90%
Similar papers in this journal
- Phase Recognition in Contrast-Enhanced CT Scans based on Deep Learning and Random Sampling 92%
- Necessity and Impact of Specialization of Large Foundation Model for Medical Segmentation Tasks 91%
- SCU-Net: A deep learning method for segmentation and quantification of breast arterial calcifications on mammograms 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.