Semantic Retrieval of Similar Radiological Images using Vision Transformers
Thakrar, A.; Jayasuriya, M.; Serapio, A.; Wu, X.; Davis, E.; Schroeder, J.; Vella, M.; Sohn, J. H.
Show abstract
BackgroundIdentifying visually and semantically similar radiological images in a database can facilitate the creation of decision support tools, teaching files, and research cohorts. Existing content-based image retrieval tools are often limited to searching by pixel-wise difference or vector distance of model predictions. Vision transformers (ViT) use attention to simultaneously take into account radiological diagnosis and visual appearance. PurposeWe aim to develop a ViT-based image retrieval framework and evaluate the algorithm on NIH Chest Radiographs (CXR) and NLST Chest CTs. Materials and MethodsThe model was trained on 112,120 CXR and 111,955 CT images. For CXR, a ViT binary classifier was trained on 4 ground truth labels (Cardiomegaly, Opacity, Emphysema, No Finding) and ensembled to produce multilabel classifications for each CXR. For CT, a regression model was trained to minimize L1 loss on the continuous ground truth labels of patient weight. The ViT image embedding layer was treated as a global image descriptor, using the L2 distance between descriptors as a similarity measure. To qualitatively evaluate the model, five radiologists performed a reader performance study with random query images (25 CT, 25 CXR). For each image, they chose the 5 most similar images from a set of 10 images (the 5 closest and 5 furthest images from the query in model space). Inter-radiologist and radiologist-model agreement statistics were calculated. ResultsThe CXR model achieved nDCG@5 of 0.73 (p<0.001) and Cardiomegaly mAP@5 of 0.76 (p<0.001) among other results. The CT model achieved nDCG of 16.85 (p<0.001). The model prediction agreed with radiologist consensus on 86% of CXR samples and 79.2% of CT samples. Inter-radiologist Fleiss Kappa of 0.51 and radiologist-consensus-to-model Cohens Kappa of 0.65 were observed. A t-SNE of the CT model latent space was generated to validate similar image clustering. ConclusionOur ViT architecture retrieved visually and semantically similar radiological images. Summary StatementThis study evaluates the efficacy of using ViT based image embeddings for CBIR tasks for CXR and CT images, finding that it performs well on visual and semantic recognition tasks. Key ResultsO_LIThe CXR model achieved nDCG@5 of 0.73 (p<0.001) and Cardiomegaly mAP@5 of 0.76 (p<0.001) among other results for CXR. C_LIO_LIThe CT model achieved nDCG of 16.85 (p<0.001). The model prediction agreed with radiologist consensus on 86% of CXR samples and 79.2% of CT samples. C_LIO_LIInter-radiologist Fleiss Kappa of 0.51 and radiologist consensus to model Cohens Kappa of 0.65 were observed. C_LI
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 96%
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 95%
- Classification of Hyper-scale Multimodal Imaging Datasets 95%
Similar papers in this journal
- Toward Understanding COVID-19 Pneumonia: A Deep-learning-based Approach for Severity Analysis and Monitoring the Disease 96%
- Effective Deep Learning Approaches for Predicting COVID-19 Outcomes from Chest Computed Tomography Volumes 96%
- COVID-Classifier: An automated machine learning model to assist in the diagnosis of COVID-19 infection in chest x-ray images 95%
Similar papers in this journal
- From Community Acquired Pneumonia to COVID-19: A Deep Learning Based Method for Quantitative Analysis of COVID-19 on thick-section CT Scans 94%
- A deep learning algorithm using CT images to screen for Corona Virus Disease (COVID-19) 94%
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 93%
Similar papers in this journal
- Enhancing Semantic Segmentation in Chest X-Ray Images through Image Preprocessing: ps-KDE for Pixel-wise Substitution by Kernel Density Estimation 97%
- ai-corona : Radiologist-Assistant Deep Learning Framework for COVID-19 Diagnosis in Chest CT Scans 96%
- Automated Detection of COVID-19 through Convolutional Neural Network using Chest x-ray images 94%
Similar papers in this journal
- Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning models 94%
- ENRICHing Medical Imaging Training Sets Enables More Efficient Machine Learning 93%
- Quantification of abdominal fat from computed tomography using deep learning and its association with electronic health records in an academic biobank 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.