Evaluation of large language model performance on the Biomedical Language Understanding and Reasoning Benchmark
Feng, H.; Ronzano, F.; LaFleur, J.; Garber, M.; de Oliveira, R.; Rough, K.; Roth, K.; Nanavati, J.; Zine El Abidine, K.; Mack, C.
Show abstract
BackgroundThe availability of increasingly powerful large language models (LLMs) has attracted substantial interest in their potential for interpreting and generating human-like text for biomedical and clinical applications. However, there are often demands for high accuracy, concerns about balancing generalizability and domain-specificity, and questions about prompting robustness when considering the adoption of LLMs for specific use cases. There also is a lack of a framework or method to help choose which LLMs (or prompting strategies) should be adopted for specific biomedical or clinical tasks. ObjectiveTo address the speculations on applying LLMs for biomedical applications, this study aims to 1) propose a framework to comprehensively evaluate and compare the performance of a range of LLMs and prompting techniques on a suite of biomedical natural language processing (NLP) tasks; 2) use the framework to benchmark several general-purpose LLMs and biomedical domain-specific LLMs. MethodsWe evaluated and compared six general-purpose LLMs (GPT-4, GPT-3.5-Turbo, Flan-T5-XXL, Llama-3-8B-Instruct, Yi-1.5-34B-Chat, and Zephyr-7B-Beta) and three healthcare-specific LLMs (Medicine-Llama3-8B, Meditron-7B, and MedLLaMA-13B) on a set of 13 datasets - referred to as the Biomedical Language Understanding and Reasoning Benchmark (BLURB) - covering six commonly needed medical natural language processing tasks: named entity recognition (NER); relation extraction (RE); population, interventions, comparators, and outcomes (PICO); sentence similarity (SS); document classification (Class.); and question-answering (QA). All models were evaluated without further training or fine-tuning. Model performance was assessed according to a range of prompting strategies (formalized as a systematic, reusable prompting framework) and relied on the standard, task-specific evaluation metrics defined by BLURB. ResultsAcross all tasks, GPT-4 outperformed other LLMs, achieving a score of 64.6 on the benchmark, though other models, such as Flan-T5-XXL and Llama-3-8B-Instruct, demonstrated competitive performance on multiple tasks. We found that general-purpose models achieved better overall scores than domain-specific models, sometimes by significant margins. We observed a substantial impact of strategically editing the prompt describing the task and a consistent improvement in performance when including examples semantically similar to the input text. Additionally, the most performant prompts for nearly half the models outperformed the previously reported best results for the PubMedQA dataset from the BLURB leaderboard. ConclusionsThese results provide evidence of the potential LLMs may have for biomedical applications and highlight the importance of robust evaluation before adopting LLMs for any specific use cases. Notably, performant open-source LLMs such as Llama-3-8B-Instruct and Flan-T5-XXL show promise for use cases where trustworthiness and data confidentiality are concerns, as these models can be hosted locally, offering better security, transparency, and explainability. Continuing to explore how these emerging technologies can be adapted for the healthcare setting, paired with human expertise, and enhanced through quality control measures will be important research to allow responsible innovation with LLMs in the biomedical area.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Annotation-preserving machine translation of English corpora to validate Dutch clinical concept extraction tools 95%
- LCD Benchmark: Long Clinical Document Benchmark on Mortality Prediction for Language Models 95%
- A Novel Question-Answering Framework for Automated Abstract Screening Using Large Language Models 95%
Similar papers in this journal
- Comparing neural language models for medical concept representation and patient trajectory prediction 93%
- Enriching Representation Learning Using 53 Million Patient Notes through Human Phenotype Ontology Embedding 92%
- The role of natural language processing in cancer care: a systematic scoping review with narrative synthesis 92%
Similar papers in this journal
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 97%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 94%
- Comparative Effectiveness of Medical Concept Embedding for Feature Engineering in Phenotyping 94%
Similar papers in this journal
- DeepPhe-CR: Natural Language Processing Software Services for Cancer Registrar Case Abstraction 93%
- Exploring Integrative Analysis using the BioMedical Evidence Graph 90%
- Actionability of Synthetic Data in a Heterogeneous and Rare Healthcare Demographic; Adolescents and Young Adults (AYAs) with Cancer 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.