Back

Medical Abbreviation Disambiguation with Large Language Models: Zero- and Few-Shot Evaluation on the MeDAL Dataset

Shafiei Rezvani Nezhad, N.; Mansouri, M.; Abolhasani, R.; Zakaria, R. A.

2025-09-17 bioinformatics
10.1101/2025.09.12.675926 bioRxiv
Show abstract

Abbreviation disambiguation is a critical challenge in processing clinical and biomedical texts, where ambiguous short forms frequently obscure meaning. In this study, we assess the zero-shot performance of large language models (LLMs) on the task of medical abbreviation disambiguation using the MeDAL dataset, a large-scale resource constructed from PubMed abstracts. Specifically, we evaluate GPT-4 and LLaMA models, prompting them with contextual information to infer the correct long-form expansion of ambiguous abbreviations without any task-specific fine-tuning. Our results demonstrate that GPT-4 substantially outperforms LLaMA across a range of ambiguous terms, indicating a significant advantage of proprietary models in zero-shot medical language understanding. These findings suggest that LLMs, even without domain-specific training, can serve as effective tools for improving readability and interpretability in biomedical NLP applications.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.