VLIB: Unveiling insights through Visual and Linguistic Integration of Biorxiv data relevant to cancer via Multimodal Large Language Model
Liu, K.; Prabhakar, V.
Show abstract
The field of cancer research has greatly benefited from the wealth of new knowledge provided by research articles and preprints on platforms like Biorxiv. This study investigates the role of scientific figures and their accompanying captions in enhancing our comprehension of cancer. Leveraging the capabilities of Multimodal Large Language Models (MLLMs), we conduct a comprehensive analysis of both visual and linguistic data in biomedical literature. Our work introduces VLIB, a substantial scientific figure-caption dataset generated from cancer biology papers on Biorxiv. After thorough preprocessing, which includes figure-caption pair extraction, sub-figure identification, and text normalization, VLIB comprises over 500,000 figures from more than 70,000 papers, each accompanied by relevant captions. We fine-tune baseline MLLMs using our VLIB dataset for downstream vision-language tasks, such as image captioning and visual question answering (VQA), to assess their performance. Our experimental results underscore the vital role played by scientific figures, including molecular structures, histological images, and data visualizations, in conjunction with their captions, in facilitating knowledge translation through MLLMs. Specifically, we achieved a ROUGE score of 0.66 for VQA and 0.68 for image captioning, as well as a BLEU score of 0.72 for VQA and 0.70 for image captioning. Furthermore, our investigation highlights the potential of MLLMs to bridge the gap between artificial intelligence and domain experts in the field of cancer biology.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- pathCLIP: Detection of Genes and Gene Relations from Biological Pathway Figures through Image-Text Contrastive Learning 94%
- Evaluating Explanations from AI Algorithms for Clinical Decision-Making: A Social Science-based Approach 93%
- SimSearch: A Human-in-the-Loop Learning Framework for Fast Detection of Regions of Interest in Microscopy Images 92%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.