ChatGPT for assessing risk of bias of randomized trials using the RoB 2.0 tool: A methods study
Pitre, T.; Jassal, T.; Talukdar, J. R.; Shahab, M.; Ling, M.; Zeraatkar, D.
Show abstract
BackgroundInternationally accepted standards for systematic reviews necessitate assessment of the risk of bias of primary studies. Assessing risk of bias, however, can be time- and resource-intensive. AI-based solutions may increase efficiency and reduce burden. ObjectiveTo evaluate the reliability of ChatGPT for performing risk of bias assessments of randomized trials using the revised risk of bias tool for randomized trials (RoB 2.0). MethodsWe sampled recently published Cochrane systematic reviews of medical interventions (up to October 2023) that included randomized controlled trials and assessed risk of bias using the Cochrane-endorsed revised risk of bias tool for randomized trials (RoB 2.0). From each eligible review, we collected data on the risk of bias assessments for the first three reported outcomes. Using ChatGPT-4, we assessed the risk of bias for the same outcomes using three different prompts: a minimal prompt including limited instructions, a maximal prompt with extensive instructions, and an optimized prompt that was designed to yield the best risk of bias judgements. The agreement between ChatGPTs assessments and those of Cochrane systematic reviewers was quantified using weighted kappa statistics. ResultsWe included 34 systematic reviews with 157 unique trials. We found the agreement between ChatGPT and systematic review authors for assessment of overall risk of bias to be 0.16 (95% CI: 0.01 to 0.3) for the maximal ChatGPT prompt, 0.17 (95% CI: 0.02 to 0.32) for the optimized prompt, and 0.11 (95% CI: -0.04 to 0.27) for the minimal prompt. For the optimized prompt, agreement ranged between 0.11 (95% CI: -0.11 to 0.33) to 0.29 (95% CI: 0.14 to 0.44) across risk of bias domains, with the lowest agreement for the deviations from the intended intervention domain and the highest agreement for the missing outcome data domain. ConclusionOur results suggest that ChatGPT and systematic reviewers only have "slight" to "fair" agreement in risk of bias judgements for randomized trials. ChatGPT is currently unable to reliably assess risk of bias of randomized trials. We advise against using ChatGPT to perform risk of bias assessments. There may be opportunities to use ChatGPT to streamline other aspects of systematic reviews, such as screening of search records or collection of data.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Exploring the potential of Claude 2 for risk of bias assessment: Using a large language model to assess randomized controlled trials with RoB 2 98%
- Evaluation of the sensitivity, accuracy and currency of the Cochrane COVID-19 Study Register for supporting rapid evidence synthesis production 96%
- Development of the Individual Participant Data (IPD) Integrity Tool for assessing the integrity of randomised trials using individual participant data 96%
Similar papers in this journal
- The use of the Registered Reports format for publication of randomized clinical trials: a cross-sectional study 96%
- Updating the PRISMA reporting guideline for network meta-analysis: a scoping review 96%
- The impact of retracted randomised controlled trials on systematic reviews and clinical practice guidelines: a meta-epidemiological study 95%
Similar papers in this journal
- Application of observational research methods to real-world studies for rare disease drugs: a scoping review protocol 94%
- Exploring scalable assessment methods for terminated trials in ClinicalTrials.gov: A cohort analysis of German and Californian trials 94%
- Investigating the use of a one-page infographic to improve recruitment and retention to the BASIL+ Randomised Controlled Trial: A Study Within a Trial (SWAT) 94%
Similar papers in this journal
- Causal Forests versus Inverse Probability of Treatment Weighting to adjust for Cluster-Level Confounding: A Parametric and Plasmode Simulation Study based on US Hosptial Electronic Health Record Data 90%
- Pregnancy pharmacoepidemiology: How often are key methodological elements reported in publications? 90%
- Bias amplification of unobserved confounding in pharmacoepidemiological studies using indication-based sampling: there is no free lunch in restricting the sample to those with a particular drug-indication 89%
Similar papers in this journal
- To Include or Not to Include? A prescription from the pharmacy on how to use active learning assisted screening in systematic reviews 93%
- Evidence-Based, Cost-Effective Interventions To Suppress The COVID-19 Pandemic: A Systematic Review 92%
- Repurposing Existing Medications for Coronavirus Disease 2019: Protocol for a Rapid and Living Systematic Review 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.