Evaluating the Reporting Quality of 21,041 Randomized Controlled Trial Articles
Srinivasan, A.; Berkowitz, J.; Kivelson, S.; Friendrich, N.; Tatonetti, N. P.
Show abstract
Incomplete reporting of a studys methods and results hinders efforts to evaluate and reproduce research findings in randomized controlled trials (RCTs), leading to potential harm. CONSORT, a set of widely endorsed RCT reporting guidelines, was designed to mitigate such harm by ensuring transparency, reproducibility and safety in RCTs. Evaluating adherence to CONSORT requires a level of language understanding that has previously precluded broad systematic assessment. However, recent advances in large language models (LLMs) now make it possible to evaluate the quality of RCTs at scale. We demonstrate that GPT-4o-mini used out of the box achieves state-of-the-art performance in evaluating RCT quality (F1 score: 0.85; precision: 0.96), with results validated by expert human annotators showing 92.24% agreement across 50 papers. Applying this tool to 21,041 open-access RCTs (1966-2024), we reveal temporal and domain trends: overall CONSORT compliance has improved substantially over time, rising from 27.3% in 1966-1990 to 56.1% in 2010-2024, yet critical methodological components remain severely underreported. Randomization procedures (9.7%), allocation concealment mechanisms (15.25%), and protocol access information (2.22%) were particularly deficient. Substantial variation exists across medical disciplines (35-63% compliance), with urology/nephrology and critical care demonstrating the highest compliance, while pharmacology showed the lowest. Trial characteristics including FDA regulation status, presence of data monitoring committees, reporting of adverse events, and mortality outcomes showed statistically significant but practically negligible differences in compliance rates. Our work provides a scalable AI framework to audit and improve RCT reporting, offering actionable insights for journals, researchers, and policymakers to enhance research integrity and clinical translation.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Improving research transparency with individualized report cards: A feasibility study in clinical trials at a large university medical center 96%
- Investigator-initiated versus industry-sponsored trials – Visibility and relevance of randomized controlled trials in clinical practice guidelines (IMPACT) 96%
- Evaluation of SURUS: a Named Entity Recognition System to Extract Knowledge from Interventional Study Records 95%
Similar papers in this journal
- Results reporting for clinical trials led by medical universities and university hospitals in the Nordic countries was often missing or delayed 95%
- The use of the Registered Reports format for publication of randomized clinical trials: a cross-sectional study 95%
- Results dissemination from clinical trials conducted at German university medical centres was delayed and incomplete 95%
Similar papers in this journal
- A modular pipeline for natural language processing-screened human abstraction of a pragmatic trial outcome from electronic health records 94%
- Dynamic methods for ongoing assessment of site-level risk in risk-based monitoring of clinical trials: a scoping review 93%
- Evidence Supporting EMA Drug Approvals (2020-2023): A Cross-Sectional Study of Trial Design and Outcomes 92%
Similar papers in this journal
- Exploring scalable assessment methods for terminated trials in ClinicalTrials.gov: A cohort analysis of German and Californian trials 97%
- Analysis of clinical trial registry entry histories using the novel R package cthist 95%
- How Informative Were Early SARS-CoV-2 Treatment and Prevention Trials? A longitudinal cohort analysis of trials registered on clinicaltrials.gov 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.