MedKLIP: Medical Knowledge Enhanced Language-Image Pre-Training
Wu, C.; Zhang, X.; Zhang, Y.; Wang, Y.; Xie, W.
Show abstract
In this paper, we consider the problem of enhancing self-supervised visual-language pre-training (VLP) with medical-specific knowledge, by exploiting the paired image-text reports from the radiological daily practice. In particular, we make the following contributions: First, unlike existing works that directly process the raw reports, we adopt a novel report filter to extract the medical entities, avoiding unnecessary complexity from language grammar and enhancing the supervision signals; Second, we propose a novel entity embedding module by querying an external knowledge description base, to exploit the rich context of additional information that the medical domain affords, and implicitly build relationships between entities in the language embedding space; Third, we propose a novel Transformer-based fusion model for spatially aligning the entity description with visual signals at the image patch level only with self-supervised learning, thus enabling the ability for spatial grounding; Fourth, we conduct thorough experiments to validate the effectiveness of our proposed architecture, and benchmark on numerous public benchmarks e.g., ChestX-ray14, RSNA Pneumonia, SIIM-ACR Pneumothorax, COVIDx CXR-2, COVID Rural, and EdemaSeverity. In both zero-shot and fine-tuning settings, our model has demonstrated strong performance compared with the former methods on disease classification and grounding.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- pathCLIP: Detection of Genes and Gene Relations from Biological Pathway Figures through Image-Text Contrastive Learning 96%
- A Transformer-Based Model Trained on Large Scale Claims Data for Prediction of Severe COVID-19 Disease Progression 93%
- Dual-Field Microvascular Segmentation: Hemodynamically-Consistent Attention Learning for Retinal Vasculature Mapping 92%
Similar papers in this journal
- Accurate recognition of colorectal cancer with semi-supervised deep learning on pathological images 95%
- Generative AI Enables Medical Image Segmentation in Ultra Low-Data Regimes 95%
- Features fusion or not: harnessing multiple pathological foundation models using Meta-Encoder for downstream tasks fine-tuning 93%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Inferring global-scale temporal latent topics from news reports to predict public health interventions for COVID-19 93%
- Generating hard-to-obtain information from easy-to-obtain information: applications in drug discovery and clinical inference 92%
- Single-Cell Multi-Modal GAN (scMMGAN) reveals spatial patterns in single-cell data from triple negative breast cancer 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.