Back

Assessing the potential of vision language models for automated phenotyping of Drosophila melanogaster

Paci, G.; Nanni, F.

2024-05-27 bioinformatics
10.1101/2024.05.27.594652 bioRxiv
Show abstract

Model organisms such as Drosophila melanogaster are extremely well suited to performing large-scale screens, which often require the assessment of phenotypes in a target tissue (e.g., wing and eye). Currently, the annotation of defects is either performed manually, which hinders throughput and reproducibility, or based on dedicated image analysis pipelines, which are tailored to detect only specific defects. Here, we assess the potential of Vision Language Models (VLMs) to automatically detect aberrant phenotypes in a dataset of Drosophila wings and provide their descriptions. We compare the performance of one the current most advanced multimodal models (GPT-4) with an open-source alternative (LLaVA). Via a thorough quantitative evaluation, we identify strong performances in the identification of aberrant wing phenotypes when providing the VLMs with just a single reference image. GPT-4 showed the best performance in terms of generating textual descriptions, being able to correctly describe complex wing phenotypes. We also provide practical advice on potential prompting strategies and highlight current limitations of these tools, especially around misclassification and generation of false information, which should be carefully taken into consideration if these tools are used as part of an image analysis pipeline.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.