Automation in Clinical Trial Statistical Programming: A Structured Review of TLF Generation, Validation Frameworks, and AI/ML Integration
Yan, J.
Show abstract
BackgroundClinical trial statistical programming is transitioning from manual, study-specific coding toward metadata-driven, automated pipelines. Although general data management transformation has been reviewed, comprehensive synthesis of statistical programming automation--particularly tables, listings, and figures (TLF) generation and validation frameworks--remains limited. This review addresses this gap through systematic evidence synthesis. MethodsWe conducted a structured literature review across PubMed, Google Scholar, arXiv, and industry conference proceedings (PharmaSUG, PHUSE, R/Pharma) from January 2020 to December 2025. We applied GRADE methodology to assess evidence quality. From 789 publications screened, 262 met inclusion criteria for synthesis. ResultsKey findings include: (1) the pharmaverse ecosystem (rtables, Tplyr, admiral) reduced TLF development time by 15-25% (GRADE [Grading of Recommendations, Assessment, Development, and Evaluation]: Low); (2) risk-based validation combined with CI/CD pipelines decreased validation effort by 30-50% (GRADE: Low); (3) metadata-driven architectures enabled 40-60% specification reuse across studies (GRADE: Very Low); (4) REDCap2SDTM reduced SDTM conversion time by 75-85% (GRADE: Moderate); (5) domain-specific large language models (LLMs) achieved 88-93% F1-scores on clinical NLP tasks (GRADE: Moderate), while general-purpose models showed 60-85% accuracy for code generation (GRADE: Very Low). Critical evidence gaps persist: only 12 of 527 validation papers (2.3%) reported quantitative outcomes, and no RCTs comparing validation approaches exist. ConclusionsClinical programming automation has reached practical maturity. However, evidence quality remains predominantly Low to Very Low. Future priorities include RCTs comparing validation approaches, standardized outcome metrics, and regulatory guidance for AI-assisted programming.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Scoping Review of Artificial Intelligence Applications in Clinical Trial Risk Assessment 95%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
- The clinician-AI interface: intended use and explainability in FDA-cleared AI devices for medical image interpretation 93%
Similar papers in this journal
- Completeness of reporting of clinical prediction models developed using supervised machine learning: A systematic review 93%
- Improving research transparency with individualized report cards: A feasibility study in clinical trials at a large university medical center 93%
- Approaches in Analyzing Predictors of Trial Failure: A Scoping Review and Meta-epidemiological study 93%
Similar papers in this journal
- Dynamic methods for ongoing assessment of site-level risk in risk-based monitoring of clinical trials: a scoping review 93%
- A modular pipeline for natural language processing-screened human abstraction of a pragmatic trial outcome from electronic health records 93%
- Evidence Supporting EMA Drug Approvals (2020-2023): A Cross-Sectional Study of Trial Design and Outcomes 91%
Similar papers in this journal
- Exploring scalable assessment methods for terminated trials in ClinicalTrials.gov: A cohort analysis of German and Californian trials 94%
- Analysis of clinical trial registry entry histories using the novel R package cthist 94%
- Clinical code sets and the problem of redundancy in code set repositories 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.