The Impact of the Temperature on Extracting Information From Clinical Trial Publications Using Large Language Models
Windisch, P.; Dennstaedt, F.; Koechli, C.; Schroeder, C.; Aebersold, D. M.; Foerster, R.; Zwahlen, D. R.
Show abstract
IntroductionThe application of natural language processing (NLP) for extracting data from biomedical research has gained momentum with the advent of large language models (LLMs). However, the effect of different LLM parameters, such as temperature settings, on biomedical text mining remains underexplored and a consensus on what settings can be considered "safe" is missing. This study evaluates the impact of temperature settings on LLM performance for a named-entity recognition and a classification task in clinical trial publications. MethodsTwo datasets were analyzed using GPT-4o and GPT-4o-mini models at nine different temperature settings (0.00-2.00). The models were used to extract the number of randomized participants and classified abstracts as randomized controlled trials (RCTs) and/or as oncology-related. Different performance metrics were calculated for each temperature setting and task. ResultsBoth models provided correctly formatted predictions for more than 98.7% of abstracts across temperatures from 0.00 to 1.50. While the number of correctly formatted predictions started to decrease afterwards with the most notable drop between temperatures 1.75 and 2.00, the other performance metrics remained largely stable. ConclusionTemperature settings at or below 1.50 yielded consistent performance across text mining tasks, with performance declines at higher settings. These findings are aligned with research on different temperature settings for other tasks, suggesting stable performance within a controlled temperature range across various NLP applications.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Is One Run Enough? Reproducibility of Flagship Large Language Models Across Temperature and Reasoning Settings in Biomedical Text Processing 96%
- Use of unstructured text in prognostic clinical prediction models: a systematic review 94%
- Analysis of Eligibility Criteria Clusters Based on Large Language Models for Clinical Trial Design 94%
Similar papers in this journal
- Optimal policy determination in sequential systemic and locoregional therapy of oropharyngeal squamous carcinomas: A patient-physician digital twin dyad with deep Q-learning for treatment selection 93%
- Empirical Sample Size Determination for Popular Classification Algorithms in Clinical Research 92%
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 92%
Similar papers in this journal
- Scalable information extraction from free text electronic health records using large language models 93%
- Evaluation of SURUS: a Named Entity Recognition System to Extract Knowledge from Interventional Study Records 93%
- Advancing data science in drug development through an innovative computational framework for data sharing and statistical analysis 91%
Similar papers in this journal
- On the predictability of postoperative complications for cancer patients: a Portuguese cohort study 92%
- MelAnalyze: Fact-Checking Melatonin claims using Large Language Models and Natural Language Inference 92%
- Automated abstraction of clinical parameters of multiple myeloma from real-world clinical notes using large language models 92%
Similar papers in this journal
- The role of natural language processing in cancer care: a systematic scoping review with narrative synthesis 93%
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 92%
- Comparing neural language models for medical concept representation and patient trajectory prediction 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.