Assessment of Bias in Clinical Trials with LLMs Using ROBUST-RCT: A Feasibility Study.
Vidor, P. R.; Casiraghi, Y.; de Souza, A. M.; Schmidt, M. I.
Show abstract
BACKGROUNDBias assessment is a crucial step in evaluating evidence from randomized controlled trials. The widely adopted Cochrane RoB 2, designed to identify these issues, is complex, resource-intensive, and unreliable. Advances in artificial intelligence (AI), particularly in the field of large language models (LLMs), now allow the automation of complex tasks. While prior investigations have focused on whether LLMs could perform assessments with RoB 2, integrating technologies does not resolve the intrinsic methodological issues of the instrument. This is the first feasibility study to evaluate the reliability of ROBUST-RCT, a novel bias assessment tool, as applied by humans and LLMs. METHODSA sample of RCTs of drug interventions was screened for eligibility. Reviewers working independently used ROBUST-RCT to assess different aspects of the studies and then reached a consensus through discussion. A chain-of-thought prompt instructed four LLMs on how to apply ROBUST-RCT. The primary analysis used Gwets AC2 coefficient and benchmarking to assess inter-rater reliability of the "judgment set", defined as the series of final assessments for the six core items in the ROBUST-RCT tool. RESULTS54 assessments of each LLM were compared to human consensus in the primary analysis. Gwets AC2 inter-rater reliability ranged from 0.46 to 0.69. With 95% confidence, three of the four tested LLMs achieved moderate or higher reliability based on probabilistic benchmarking. A secondary analysis also found a Fleiss Kappa of 0.49 (95% CI: 0.30 - 0.60) between human reviewers before consensus, numerically higher than the values reported in prior literature about RoB 2. CONCLUSIONLarge Language Models (LLMs) can effectively perform risk-of-bias assessments using the ROBUST-RCT tool, enabling their integration into future systematic review workflows aiming for enhanced objectivity and efficiency.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The use of the Registered Reports format for publication of randomized clinical trials: a cross-sectional study 97%
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 97%
- The impact of retracted randomised controlled trials on systematic reviews and clinical practice guidelines: a meta-epidemiological study 96%
Similar papers in this journal
- Exploring the potential of Claude 2 for risk of bias assessment: Using a large language model to assess randomized controlled trials with RoB 2 98%
- Evaluation of the sensitivity, accuracy and currency of the Cochrane COVID-19 Study Register for supporting rapid evidence synthesis production 96%
- Development of the Individual Participant Data (IPD) Integrity Tool for assessing the integrity of randomised trials using individual participant data 96%
Similar papers in this journal
- Investigator-initiated versus industry-sponsored trials – Visibility and relevance of randomized controlled trials in clinical practice guidelines (IMPACT) 96%
- Improving research transparency with individualized report cards: A feasibility study in clinical trials at a large university medical center 95%
- Does pre-notification increase questionnaire response rates: a nested randomised control trial 94%
Similar papers in this journal
- Exploring scalable assessment methods for terminated trials in ClinicalTrials.gov: A cohort analysis of German and Californian trials 95%
- COVID-19-related research data availability and quality according to the FAIR principles: A meta-research study 93%
- Modelling the impact of behavioural interventions during pandemics: A systematic review 93%
Similar papers in this journal
- GPT for RCTs?: Using AI to measure adherence to reporting guidelines 96%
- Protocol for the development of a tool (INSPECT-SR) to identify problematic randomised controlled trials in systematic reviews of health interventions 96%
- Reporting of Retrospective Registration in Clinical Trial Publications 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.