More Signal vs. More Noise - Comparing Full Text and Abstract as Inputs for Large Language Model-based Classification of Oncology Trial Eligibility Criteria
Weyrich, J.; Dennstaedt, F.; Foerster, R.; Schroeder, C.; Aebersold, D. M.; Zwahlen, D. R.; Windisch, P.
Show abstract
PurposeLarge language models (LLMs) offer significant potential for automating the classification of clinical trials by eligibility criteria. However, a critical question remains regarding the optimal input data: while abstracts provide a condensed, high-density signal, full-text articles contain a much higher volume of information. It remains unclear whether the additional signal found in full texts improves classification performance or if the accompanying noise (in the form of thousands of words irrelevant to the question at hand in a complete manuscript) negatively affects the models reasoning capabilities. MethodsGPT-5 was applied to classify 200 randomized controlled oncology trials from high-impact medical journals, labeling them whether patients with localized and/or metastatic disease were eligible for inclusion. Each trial was classified twice - once using only the abstract and once using the full text - and GPT-5s outputs were compared with the ground-truth labels established by manual annotation. Performance was assessed by calculating and comparing accuracy, precision, recall, and F1 score, and the McNemar test was used to assess the statistical significance of the differences between the two input formats. ResultsFor identifying trials including patients with localized disease, GPT-5 achieved an accuracy of 86% (95% CI: 81% - 91%; F1 = 0.90) when using abstracts and 92% (95% CI: 88% - 95%; F1 = 0.92) when using full texts (p = 0.027). Performance for detecting trials, which include patients with metastatic disease, was comparably high, with accuracies of 99% (95% CI: 99% - 100%; F1 = 1.00) based on abstracts and 98% (95% CI: 97% - 100%; F1 = 0.99) based on full texts. Overall accuracy for assigning combined labels per trial increased from 86% (95% CI: 81% - 91%) using abstracts to 92% (95% CI: 88% - 95%) using full texts (p = 0.027). ConclusionProviding full-text articles to GPT-5 significantly improved the classification of trial eligibility criteria. These findings suggest that, for this task, the benefit of the additional signal contained within the full text outweighed the potential for performance degradation caused by increased noise. Utilizing full-text analysis appears particularly valuable for extracting specific eligibility criteria in oncology that are frequently omitted or not explicitly described within the abstract.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Towards Predicting 30-Day Readmission among Oncology Patients: Identifying Timely and Actionable Risk Factors 94%
- Large language models to help appeal denied radiotherapy services 94%
- NCT Precision Oncology Thesaurus Drugs – a Curated Database for Drugs, Drug Classes, and Drug Targets in Precision Cancer Medicine 92%
Similar papers in this journal
- Is One Run Enough? Reproducibility of Flagship Large Language Models Across Temperature and Reasoning Settings in Biomedical Text Processing 95%
- A Web-based Tool for Automatically linking Clinical Trials to their Publications 95%
- Use of unstructured text in prognostic clinical prediction models: a systematic review 94%
Similar papers in this journal
- Investigator-initiated versus industry-sponsored trials – Visibility and relevance of randomized controlled trials in clinical practice guidelines (IMPACT) 94%
- Completeness of reporting of clinical prediction models developed using supervised machine learning: A systematic review 94%
- Evaluation of SURUS: a Named Entity Recognition System to Extract Knowledge from Interventional Study Records 93%
Similar papers in this journal
- Strength of Statistical Evidence for the Efficacy of Cancer Drugs: A Bayesian Re-Analysis of Trials Supporting FDA Approval 95%
- Results dissemination from clinical trials conducted at German university medical centres was delayed and incomplete 93%
- Results reporting for clinical trials led by medical universities and university hospitals in the Nordic countries was often missing or delayed 93%
Similar papers in this journal
- Optimal policy determination in sequential systemic and locoregional therapy of oropharyngeal squamous carcinomas: A patient-physician digital twin dyad with deep Q-learning for treatment selection 93%
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 92%
- Empirical Sample Size Determination for Popular Classification Algorithms in Clinical Research 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.