From Protocol to Analysis Plan: Development and Validation of a Large Language Model Pipeline for Statistical Analysis Plan Generation using Artificial Intelligence (SAPAI)
Jafari, H.; Chu, P.; Lange, M.; Maher, F.; Glen, C.; Pearson, O. J.; Burges, C.; Martyn, M.; Cross, S.; Carter, B.; Emsley, R.; Forbes, G.
Show abstract
Background: Statistical Analysis Plans (SAPs) are essential for trial transparency and credibility but are resource-intensive to produce. While Large Language Models (LLMs) have shown promise in drafting protocols, their ability to generate high-quality, protocol-compliant SAPs remains untested against current content guidance. This study developed and validated an LLM-based pipeline for drafting SAPs from clinical trial protocols. Methods: We developed a structured, section-by-section prompting pipeline aligned with standard SAP guidance. We applied this pipeline to nine clinical trial protocols using three leading LLMs: OpenAI GPT-5, Anthropic Claude Sonnet 4, and Google Gemini 2.5 Pro. The resulting 27 SAPs were evaluated against a 46-item quality checklist derived from the published SAP guidelines. Items were double-scored by independent trial statisticians on a 0 to 3 scale for accuracy. We compared performance across LLMs and between item types (descriptive vs. statistical reasoning) using mixed-effects logistic regression. Results: Across 9 trials, the models produced SAP drafts with high overall accuracy (77% to 78%), with no difference in performance between the three LLMs (p=0.79) but varied by content type (p < 0.001). All models performed well on descriptive items (e.g., administrative details, trial design), with lower accuracy for items requiring statistical reasoning (e.g., modelling strategies, sensitivity analyses). Accuracy for statistical items ranged from 67% to 72%, whereas descriptive items achieved 81% to 83% accuracy. Qualitatively, models were prone to specific failure modes in complex sections, such as omitting necessary details for secondary outcome models or hallucinating sensitivity analyses. Discussion: Current LLMs can effectively draft portions of SAPs, offering the potential for substantial time savings in trial documentation. However, a human-in-the-loop approach remains mandatory; while models demonstrate strong capability in producing descriptive content, their independent application to complex statistical methodology design still requires further methodological development and training. Future work should explore advanced prompt engineering, such as retrieval-augmented generation or agentic workflows, to improve reasoning capabilities.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Controlled evaLuation of Angiotensin Receptor Blockers for COVID-19 respIraTorY disease (CLARITY): Statistical analysis plan for a randomised controlled Bayesian adaptive sample size trial 96%
- Trials that turn from retrospectively registered to prospectively registered: A cohort study of ‘retroactively prospective’ clinical trial registration using history data 95%
- Development of a consensus extension of the estimands framework for cluster randomised trials (CRT-estimands): results from an international Delphi study 95%
Similar papers in this journal
- A modular pipeline for natural language processing-screened human abstraction of a pragmatic trial outcome from electronic health records 94%
- Use of estimands in cluster randomised trials: a review 94%
- Dynamic methods for ongoing assessment of site-level risk in risk-based monitoring of clinical trials: a scoping review 93%
Similar papers in this journal
- Approaches in Analyzing Predictors of Trial Failure: A Scoping Review and Meta-epidemiological study 95%
- Investigator-initiated versus industry-sponsored trials – Visibility and relevance of randomized controlled trials in clinical practice guidelines (IMPACT) 95%
- Improving research transparency with individualized report cards: A feasibility study in clinical trials at a large university medical center 94%
Similar papers in this journal
- Exploring scalable assessment methods for terminated trials in ClinicalTrials.gov: A cohort analysis of German and Californian trials 96%
- Analysis of clinical trial registry entry histories using the novel R package cthist 96%
- Investigating the use of a one-page infographic to improve recruitment and retention to the BASIL+ Randomised Controlled Trial: A Study Within a Trial (SWAT) 95%
Similar papers in this journal
- The use of the Registered Reports format for publication of randomized clinical trials: a cross-sectional study 96%
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 95%
- Results reporting for clinical trials led by medical universities and university hospitals in the Nordic countries was often missing or delayed 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.