Large Language Models for Supporting Clear Writing and Detecting Spin in Randomized Controlled Trials in Oncology
Koechli, C.; Dennstaedt, F.; Schroeder, C.; Aebersold, D. M.; Foerster, R.; Zwahlen, D. R.; Windisch, P.
Show abstract
ImportanceAccurate interpretation of randomized controlled trial (RCT) results is essential for guiding clinical practice in oncology. Reporting "spin" can misrepresent treatment efficacy, potentially leading to suboptimal clinical decisions. Standardized methods could help to detect misleading reporting. ObjectiveTo determine whether large language models (LLMs) can accurately classify oncology RCTs as positive or negative based on primary endpoint achievement when provided with different sections of the trial report, thereby assessing their utility in identifying potential spin in conclusions. DesignMethodological evaluation using LLMs and human annotations of previously published clinical trials. SettingRandom sample of RCTs from seven major medical journals published between 2005 and 2023. Participants250 two-arm, single primary endpoint oncology RCT reports were randomly selected from the specified journals and publication years. Exposure(s)Human annotators independently classified trials based on primary endpoint results before LLM evaluation. Three commercial LLMs (GPT-3.5 Turbo, GPT-4o, and o1) classified trials based on four different text inputs: 1) conclusion only, 2) methods and conclusion, 3) methods, results, and conclusion, or 4) title and full abstract. Main Outcome(s) and Measure(s)Performance of LLMs in classifying trials as positive or negative, primarily measured using the F1 score. ResultsThe analysis included 250 RCT reports; based on human annotation, 146 (58.4%) were positive and 104 (41.6%) were negative. o1 demonstrated the highest performance across all input conditions, achieving F1 scores of 0.932 (conclusion only), 0.960 (methods and conclusion), 0.980 (methods, results, and conclusion), and 0.970 (title and full abstract). Analysis of trials incorrectly classified as positive by the LLM when using only the conclusion revealed patterns such as absence of primary endpoint data, emphasis on secondary or subgroup findings, or unclear endpoint distinctions within the conclusion. Conclusions and RelevanceLLMs can accurately classify oncology RCT outcomes. Discrepancies between classifications based on conclusions versus more complete text indicate potential spin. This approach could serve as a valuable supplementary tool for researchers, reviewers, and editors to enhance transparency and critical appraisal of oncology trial reporting, though further validation is required, especially for trials with more complex designs. Key PointsO_ST_ABSQuestionC_ST_ABSCan large language models (LLMs) accurately classify oncology randomized controlled trials (RCTs) as positive or negative based on primary endpoint achievement, and can this help identify potential "spin" in trial conclusions? FindingsIn this methodological evaluation of 250 oncology RCTs, the o1 LLM achieved high accuracy in classifying trials based on the title and full abstract, outperforming classifications based on the conclusion alone. Discrepancies between classifications using conclusions versus more complete text often indicated patterns that could be considered as spin. MeaningLLMs show promise as a supplementary tool to detect potential spin by identifying inconsistencies between conclusions and overall results.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Completeness of reporting of clinical prediction models developed using supervised machine learning: A systematic review 94%
- Investigator-initiated versus industry-sponsored trials – Visibility and relevance of randomized controlled trials in clinical practice guidelines (IMPACT) 94%
- Predictive approaches to heterogeneous treatment effects: a systematic review 93%
Similar papers in this journal
- Is One Run Enough? Reproducibility of Flagship Large Language Models Across Temperature and Reasoning Settings in Biomedical Text Processing 94%
- A Web-based Tool for Automatically linking Clinical Trials to their Publications 93%
- Use of unstructured text in prognostic clinical prediction models: a systematic review 93%
Similar papers in this journal
- Strength of Statistical Evidence for the Efficacy of Cancer Drugs: A Bayesian Re-Analysis of Trials Supporting FDA Approval 94%
- The use of the Registered Reports format for publication of randomized clinical trials: a cross-sectional study 94%
- The impact of retracted randomised controlled trials on systematic reviews and clinical practice guidelines: a meta-epidemiological study 93%
Similar papers in this journal
- Surgical Resection, Radiotherapy, And Percutaneous Thermal Ablation for Treatment of Stage 1 Non-Small Cell Lung Cancer: A Systematic Review and Network Meta-Analysis 94%
- Reproducibility and transparency characteristics of oncology research evidence 94%
- GPT for RCTs?: Using AI to measure adherence to reporting guidelines 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.