Back

Fully automatic summarization of radiology reports using natural language processing with language models.

Nishio, M.; Matsunaga, T.; Matsuo, H.; Nogami, M.; Kurata, Y.; Fujimoto, K.; Sugiyama, O.; Akashi, T.; Aoki, S.; Murakami, T.

2023-12-02 radiology and imaging
10.1101/2023.12.01.23299267 medRxiv
Show abstract

Natural language processing using language models has yielded promising results in various fields. The use of language models may help improve the workflow of radiologists. This retrospective study aimed to construct and evaluate language models for the automatic summarization of radiology reports. Two datasets of radiology reports were used: MIMIC-CXR and the Japan Medical Image Database (JMID). MIMIC-CXR is an open dataset comprising chest radiograph reports. JMID is a large dataset of CT and MRI reports comprising reports from 10 academic medical centers in Japan. A total of 128,032 and 1,101,271 reports from the MIMIC-CXR and JMID, respectively, were included in this study. Four Text-to-Text Transfer Transformer (T5) models were constructed. Recall-Oriented Understudy for Gisting Evaluation (ROUGE), a quantitative metric, was used to evaluate the quality of text summarized from 19,205 and 58,043 test sets from MIMIC-CXR and JMID, respectively. The Wilcoxon signed-rank test was utilized to evaluate the differences among the ROUGE values of the four T5 models. In addition, subsets of automatically summarized text in the test sets were manually evaluated by two radiologists. Based on the Wilcoxon signed-rank test, the best T5 models were selected for the automatic summarization. The quantitative metrics of the best T5 models were as follows: ROUGE-1 = 57.75 {+/-} 30.99, ROUGE-2 = 49.96 {+/-} 35.36, and ROUGE-L = 54.07 {+/-} 32.48 in MIMIC-CXR; ROUGE-1 = 50.00 {+/-} 29.24, ROUGE-2 = 39.66 {+/-} 30.21, and ROUGE-L = 47.87 {+/-} 29.44 in JMID. The radiologists evaluations revealed that 86% (86/100) and 85% (85/100) of the texts automatically summarized from MIMIC-CXR and JMID, respectively, were clinically useful. The T5 models constructed in this study were capable of automatic summarization of radiology reports. The radiologists evaluations revealed that most of the automatically summarized texts were clinically valuable.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.