Large Language Models for Detecting CONSORT Guideline Compliance in Published Randomized Clinical Trials: A Cross-Sectional Evaluation Study
Tsybulnik, D. Y.; Gillette, J. J.; Heston, T. F.
Show abstract
BackgroundPeer review processes may inadequately assess compliance with established reporting guidelines such as the Consolidated Standards of Reporting Trials (CONSORT) criteria. Large language models (LLMs) demonstrate potential for systematic manuscript evaluation; however, their accuracy in detecting adherence to CONSORT guidelines in published clinical trials remains unexplored. MethodsThis cross-sectional study evaluated the compliance of 20 randomized controlled trials published between 2015 and 2024 from immunology journals, identified through PubMed, with the CONSORT 2010 guidelines. Three large language models (ChatGPT-4o, Gemini 2.5 Pro, and Claude Sonnet 4) independently assessed compliance across 37 CONSORT subpoints. The primary endpoint was the mean CONSORT compliance percentage. Secondary endpoints included the proportion of articles meeting a 90% compliance threshold and agreement between LLM assessments. Statistical analysis employed repeated measures ANOVA with post-hoc pairwise comparisons ( = 0.05). ResultsMean CONSORT compliance rates were: ChatGPT-4o 81% (95% CI: 77-85%), Claude Sonnet 4 68% (95% CI: 61-75%), and Gemini 2.5 Pro 55% (95% CI: 48-62%). Overall compliance across all LLMs was 68% (95% CI: 64-72%). Using a 90% compliance threshold as a quality benchmark, ChatGPT-4o identified 25% of papers (5/20), Claude Sonnet 4 identified 5% (1/20), and Gemini 2.5 Pro identified none (0/20) as meeting this standard. Repeated-measures ANOVA demonstrated significant differences in LLM performance (F2,38 = 40.79, p < 0.001, partial 2 = 0.682). All pairwise comparisons between models were statistically significant (p [≤] 0.002). ConclusionsLarge language models detected CONSORT compliance deficiencies in published randomized trials, aligning with previously reported rates of 60-70%, which validates their accuracy in identifying persistent reporting quality issues. The substantial variation between LLM assessments indicates the need for standardized evaluation protocols. These findings support the potential utility of LLM-assisted manuscript evaluation to improve adherence to established reporting guidelines.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- COVID-19-related research data availability and quality according to the FAIR principles: A meta-research study 96%
- Transparency in peer review: Exploring the content and tone of reviewers' confidential comments to editors 95%
- Application of observational research methods to real-world studies for rare disease drugs: a scoping review protocol 94%
Similar papers in this journal
Similar papers in this journal
- Investigator-initiated versus industry-sponsored trials – Visibility and relevance of randomized controlled trials in clinical practice guidelines (IMPACT) 94%
- Development, validation, and usage of metrics to evaluate clinical research hypothesis quality 93%
- Open Science Saves Lives: Lessons from the COVID-19 Pandemic 93%
Similar papers in this journal
Similar papers in this journal
- The use of the Registered Reports format for publication of randomized clinical trials: a cross-sectional study 96%
- Results reporting for clinical trials led by medical universities and university hospitals in the Nordic countries was often missing or delayed 94%
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.