Back

Clinical Trials

SAGE Publications

All preprints, ranked by how well they match Clinical Trials's content profile, based on 11 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Operational complexity predicts selective non-dissemination within pharmaceutical sponsor portfolios: a retrospective cohort study

Sayed, A. M.; Huan, P. T.; Nguyen, T. K.; Fathy, E.; Aziz, T.; Tho, D. V.; Huy, N. T.

2026-05-06 health policy 10.64898/2026.05.05.26352331 medRxiv
Top 0.1%
34.6%
Show abstract

BackgroundIncomplete dissemination of clinical trial results remains an important challenge for research transparency and evidence synthesis. Although prior studies have quantified the overall extent of non-dissemination, less is known about whether trial characteristics observable at registration are associated with subsequent dissemination within sponsor portfolios. Methods and findingsWe conducted a retrospective cohort study of 17,537 completed interventional clinical trials registered on ClinicalTrials.gov between 2007 and 2024 across the 20 largest global pharmaceutical companies. We developed the Operational Complexity Index (OCI), a composite measure derived from planned enrollment, facility count, and geographic scope, and examined its association with trial dissemination using multivariable logistic regression and time-to-event analyses. Higher OCI was associated with greater odds of dissemination (adjusted odds ratio [aOR] = 2.40, 95% CI 2.23-2.60; p < 0.001), with dissemination increasing from 47% in the lowest OCI decile to 95% in the highest. Higher operational complexity was also associated with earlier dissemination; over a 1,095-day horizon, high-OCI trials were disseminated a mean of 310.88 days earlier than low-OCI trials (RMST difference, 310.88 days; 95% CI 300.59-320.96; p < 0.001). This pattern was observed across sponsors, clinical phases, and therapeutic areas. In predictive analyses using registration-time variables, the structural model achieved a cross- validated AUC of 0.816 and a holdout AUC of 0.814, whereas the full model, including sponsor identity, achieved a cross-validated AUC of 0.858 and a holdout AUC of 0.857. Using benchmark phase-based costing assumptions, the 5,019 non-disseminated trials corresponded to an estimated US$10.94-15.26 billion in sunk research investment. ConclusionsAmong trials conducted by the 20 largest pharmaceutical sponsors, greater operational complexity at registration was associated with a higher likelihood of dissemination and earlier dissemination. These findings suggest that aggregate sponsor-level transparency metrics may mask important heterogeneity within sponsor portfolios. Future work should assess whether registration-time trial characteristics can help identify trial subgroups at higher risk of non-dissemination. AUTHOR SUMMARYO_ST_ABSWhy was this study done?C_ST_ABSO_LIIncomplete dissemination of clinical trial results reduces the completeness of the medical evidence base and the public value of research participation. C_LIO_LIPrevious studies have described overall rates of trial non-dissemination, but less is known about whether dissemination varies systematically across different types of trials within sponsor portfolios. C_LIO_LIWe examined whether trial characteristics available at registration were associated with later dissemination of results among large pharmaceutical sponsors. C_LI What did the researchers do and find?O_LIWe analyzed 17,537 completed interventional clinical trials sponsored by the 20 largest pharmaceutical companies and registered on ClinicalTrials.gov between 2007 and 2024. C_LIO_LIWe developed an Operational Complexity Index (OCI) based on planned enrollment, number of facilities, and geographic scope to measure trial operational scale at registration. C_LIO_LIHigher OCI was associated with a greater likelihood of dissemination and earlier dissemination. Dissemination ranged from 47% in the lowest OCI decile to 95% in the highest. C_LIO_LIThis pattern was observed across sponsor portfolios, clinical phases, and therapeutic areas, with an average within-sponsor dissemination gap of 40 percentage points between lower- and higher-complexity trials. C_LIO_LIIn manual validation of 344 sampled trials, the automated dissemination-classification pipeline achieved 92.1% accuracy. C_LIO_LIUsing benchmark phase-based costing assumptions, the 5,019 non-disseminated trials corresponded to an estimated US$10.9-15.3 billion in sunk research investment. C_LI What do these findings mean?O_LIDissemination was not uniform across trial types within sponsor portfolios; trials with lower operational complexity were less likely to be disseminated than trials with higher operational complexity. C_LIO_LIAggregate sponsor-level transparency measures may therefore miss important differences within portfolios. C_LIO_LIRegistration-time trial characteristics showed predictive signal for non-dissemination, but whether such information could support monitoring strategies would require prospective validation. C_LIO_LIMore complete dissemination of trial results would strengthen the scientific record and improve the public value of clinical research. C_LI

2
Assessing the Compliance and Timeliness of Results Reporting for Clinical Trials on Antimicrobial Agents

Curtin, M.; Wiltshire, A.; Nilsonne, G.; Siebert, M.

2025-12-31 health policy 10.64898/2025.12.23.25342922 medRxiv
Top 0.1%
26.4%
Show abstract

ObjectiveAntimicrobial resistance (AMR) is an urgent global health threat, resulting in more than 5 million deaths globally in 2019. Timely and complete antimicrobial agent (AMA) clinical trial results reporting is essential to evaluate the safety and efficacy of investigational therapies. The Food and Drug Administration Amendments Act (FDAAA) of 2007 mandated results reporting for applicable clinical trials to ClinicalTrials.gov. After nearly ten years of underreporting, the HHS issued the Final Rule, requiring a designated responsible party to submit results to ClinicalTrials.gov and clarifying applicable clinical trial (ACT) criteria. ACTs and probable ACTs (pACTs) are interventional studies regulated by the FDA with at least one site based in the United States. However, pACTs were initiated prior to January 2017, when the Final Rule came into effect. This study investigates the compliance and timeliness of results reporting of ACTs and probable ACTs (pACTs) for AMAs. DesignWe extracted data from ClinicalTrials.gov for trials involving AMAs with primary completion dates between May 1, 2013, and May 1, 2023. We analyzed the time from primary completion to results reporting and estimated the hazard ratio to compare timeliness between ACTs and pACTs. Additionally, we assessed delays in reporting across different study types and funding sources. ResultsOur search resulted in 2629 NCT records. After exclusion of ineligible trials, we included 2525 trials. We found 1769 pACTs (70.1%; 95% CI, 69.3%-72.9%) and 756 ACTs (29.9%; 95% CI, 28.2%-31.8%). Among the 2525 eligible trials, 2249 trials (89.1%; 95% CI, 87.8%-90.2%) were reported on ClinicalTrials.gov or in journal publications. Overall, 81.3% (95% CI, 79.7%-82.3%) of trials were reported late or missing (75.0% of ACTs vs 83.6% of pACTs). ACTs were more likely to report results earlier than pACTs, with a hazard ratio of 1.4 (95% CI, 1.3-1.5). ConclusionsACTs demonstrated greater reporting compliance and shorter delays in the reporting of overdue results. While this analysis provides initial insights, limitations related to timeline and sample scope suggest that broader investigations are needed to fully evaluate the impact of the Final Rule.

3
Preregistration and Credibility of Clinical Trials

Decker, C.; Ottaviani, M.

2023-05-23 health policy 10.1101/2023.05.22.23290326 medRxiv
Top 0.1%
26.3%
Show abstract

Preregistration at public research registries is considered a promising solution to the credibility crisis in science, but empirical evidence of its actual benefit is limited. Guaranteeing research integrity is especially vital in clinical research, where human lives are at stake and investigators might suffer from financial pressure. This paper analyzes the distribution of p-values from pre-approval drug trials reported to ClinicalTrials.gov, the largest registry for research studies in human volunteers, conditional on the preregistration status. The z-score density of non-preregistered trials displays a significant upward discontinuity at the salient 5% threshold for statistical significance, indicative of p-hacking or selective reporting. The density of preregistered trials appears smooth at this threshold. With caliper tests, we establish that these differences between preregistered and non-preregistered trials are robust when conditioning on sponsor fixed effects and other design features commonly indicative of research integrity, such as blinding and data monitoring committees. Our results suggest that preregistration is a credible signal for the integrity of clinical trials, as far as it can be assessed with the currently available methods to detect p-hacking.

4
PED-X-Bench: A Benchmark of Adult-to-Pediatric Extrapolation Decisions in FDA Drug Labels

Srinivasan, A.; Berkowitz, J. S.; Friedrich, N.; Tsang, K.; Kuchi, A.; Acitores Cortina, J. M.; Zietz, M.; Czarny, R.; Liu, H.; Tatonetti, N. P.

2025-05-23 health systems and quality improvement 10.1101/2025.05.22.25328187 medRxiv
Top 0.1%
23.2%
Show abstract

Pediatric trials are ethically and logistically difficult, so the U.S. FDA often extrapolates adult data to children when justified. Yet no public resource systematically documents these decisions. We present PED-X-Bench, the first dataset and benchmark that encodes FDA pediatric-extrapolation outcomes as a four-way classification task (Full, Partial, None, Unlabeled). PED-X-Bench contains 737 FDA drug-label sections ({approx} 1 M words of source text) for approvals issued 2007-2024 across all therapeutic areas. A two-stage o3-mini prompting pipeline mined full FDA label text; nine domain reviewers then adjudicated a stratified sample of 135 labels yielding an accuracy F1 of 0.74 and 0.63 respectively (inter-annotator {kappa} = 0.678) and spot-checking the remainder. For every drug we release the ground-truth label, concise efficacy and pharmacokinetic/safety summaries, and harmonized study metadata. To showcase utility we release two baseline models: (i) a logistic-regression classifier that uses structured metadata from FDAs pediatric-labeling dataset, and (ii) a fine-tuned BigBird BERT that ingests full label text. Both base-lines perform modestly, leaving ample headroom for future work. PED-X-Bench enables research on pediatric drug development, clinical NLP and drug safety; dataset card and code are made available here: github.com/tatonetti-lab/PedXBench huggingface.co/datasets/apoorvasrinivasan/Ped-X-Bench

5
MedDRA Adoption and Adverse Event Reporting Quality in Gastrointestinal and Abdominal Surgery Randomized Controlled Trials: A Cross-Sectional Analysis

Camasso, N.; Kirby, K.; Calvert, N.; Stroup, J.; Langerman, R.; Vassar, M.

2026-02-05 health systems and quality improvement 10.64898/2026.02.04.26345608 medRxiv
Top 0.1%
18.8%
Show abstract

IntroductionAdverse event (AE) reporting transparency is essential for evidence-based surgical practice, yet substantial reporting gaps persist despite Consolidated Standards for Reporting Trials (CONSORT) Harms guidance. The Medical Dictionary for Regulatory Activities (MedDRA) provides standardized terminology for AE classification, but its association with AE reporting quality remains unexplored. ObjectivesThe purpose of this study was to establish the frequency of Medical Dictionary for Regulatory Activities (MedDRA) utilization in gastrointestinal and abdominal surgical trials, identify predictors of its adoption, and quantify the association between MedDRA use and adverse event reporting quality as measured by Completeness scores, registry-publication Concordance, and overall Transparency indices. DesignCross sectional analysis of matched randomized controlled trial registry-publication pairs. Participants116 gastrointestinal and abdominal surgery randomized controlled trials registered on ClinicalTrials.gov with results posted between September 2009 and December 2024 and an associated peer-reviewed publication. Primary and Secondary Outcome MeasuresPrimary outcomes were differences in AE reporting quality between MedDRA-documenting and non-documenting trials, measured using Harms Reporting Completeness score (0-8), Concordance score (0-7), and Harms Transparency Index (0-15). Secondary outcomes included prevalence of MedDRA adoption and predictors of MedDRA documentation via univariable logistic regression. ResultsAmong 116 included trials, only 22 (18.8%) explicitly documented MedDRA use. Industry-funded trials (OR=29.32, 95% CI=8.94-118.50, p<0.001) and those with at least one U.S. site (OR=4.59, 95% CI=1.22-30.02, p=0.050) demonstrated significantly higher rates of MedDRA adoption. Trials documenting MedDRA use demonstrated significantly improved reporting across all three score parameters: Completeness score (p<0.001), Concordance score (p=0.002), and Transparency Index (p<0.001). MedDRA use was also associated with lower rates of registry-publication discordance across key safety metrics: serious adverse event (SAE) participant count registry-publication discordance was 59.1% in MedDRA documenting trials and 85.1% in non-MedDRA trials; mortality reporting discordance was 60.0% in MedDRA trials and 82.1% in non-MedDRA trials. ConclusionDespite strong association with improved AE reporting completeness and registry-publication concordance, MedDRA adoption in gastrointestinal and abdominal surgical trials remains below 20%, concentrated among industry-funded studies. The predominance of unstandardized terminology and free-text strategies promotes reporting inadequacies that complicate evidence synthesis and undermine evidence-based surgical practice. Journals, funding agencies, academic institutions, and researchers should prioritize the adoption of standardized AE terminology to enhance transparency and improve surgical research. Trial RegistrationPROSPERO CRD420251081191. Strengths and Limitations of this StudyO_LIThis is the first study to quantify the association between MedDRA use and adverse event reporting quality in surgical trials C_LIO_LIDual independent screening and extraction with pre-registered protocol minimizes bias and enhances reproducibility C_LIO_LIAnalysis limited to gastrointestinal and abdominal surgery; generalizability to other surgical subspecialties remains uncertain C_LIO_LIRequired explicit MedDRA documentation; trials using MedDRA without disclosure would be misclassified as non-users C_LIO_LIConcordance assessment examined numerical agreement without evaluating clinical significance of discrepancies C_LI

6
A Measurement-Based Care Strategy for Buprenorphine-Naloxone Treatment (Bup-MBC): Development of an EHR-Integrated Intervention

Reese, T.; Audet, C.; Ancker, J.; Wright, A.; Marcovitz, D.; Kast, K. A.; Bridges, J.; Tindle, H.; Shah, M.; von Horn, A.; Matheny, M. E.

2026-09-01 addiction medicine 10.64898/2026.08.27.26361539 medRxiv
Top 0.1%
18.6%
Show abstract

Introduction: Risk of recurrent opioid use during buprenorphine-naloxone (bup-nx) treatment is dynamic and remains elevated after initiation, with vulnerability shaped in part by treatment intensity and gaps between visits, yet routine outpatient care relies on episodic encounters and retrospective data. This mismatch can delay recognition of emerging instability and limit timely treatment adjustments. This paper reports the development and specification of an intervention strategy to address this mismatch. Methods: We used a structured, multi-phase design process to specify and configure a measurement-based care (MBC) strategy for bup-nx treatment (Bup-MBC) in outpatient addiction clinics through three phases: (1) a systematic review of patient-reported outcome measures (PROMs) for substance use treatment; (2) a qualitative needs assessment using the Theoretical Domains Framework and COM-B (Capability, Opportunity, Motivation-Behavior) model to identify gaps in risk monitoring, agency, and trust; and (3) iterative co-design with multidisciplinary clinicians to refine workflow fit and trust-preserving use of data. Patients informed item and feedback content during the needs assessment but did not participate in the co-design cycles. Results: Bup-MBC integrates (1) brief between-visit PROMs (e.g., withdrawal, craving, adherence); (2) immediate non-punitive patient feedback; (3) clinician-facing summaries and non-directive prompts in the electronic health record (EHR); and (4) an opt-in between-visit outreach pathway with predefined safety triggers, all configured within existing EHR and patient portal infrastructure. It targets patient and clinician capability to recognize changes in risk, opportunity for action through structured monitoring and visit preparation, and trust and agency through non-punitive communication, without adding substantial burden. The full measure set, severity bands, and question-to-action map are provided as supplementary material. Key trade-offs included prioritizing single-item measures for feasibility, balancing opt-in outreach with safety overrides, and assuming routine clinician use of summaries. Conclusion: This development study specifies an EHR-integrated MBC strategy for outpatient bup-nx treatment. As single-center design work with co-design limited to clinicians and delivery contingent on portal or text-message access, its outputs are hypotheses about mechanism and fit rather than demonstrated effects. Feasibility studies are needed to evaluate uptake, acceptability, workflow fit, and effects on treatment.

7
Divergent views on drivers of early phase clinical trial participation among ethnically diverse potential trial participants in the United Kingdom: A Mixed Methods Study

Hortelano, P. A.; Morton, N.; Wicks, P.; Young, M.; Burdell, R.; Richards, D.

2024-04-05 health systems and quality improvement 10.1101/2024.04.04.24305355 medRxiv
Top 0.1%
18.0%
Show abstract

BackgroundNovel therapeutics should always be tested in a sample representative of the population in need of treatment. Initial efforts of drug development take place in early phase trials (phase-I and -II), setting the direction for late-stage studies (phase-III and -IV). However, study samples in early phase trials typically fail to recruit Black, Asian and minority ethnic groups, which might produce results which dont generalise to a broader population in later trials, and ultimately, clinical practice. Focusing on early phase clinical trials the present study (1) explored the barriers and incentives that determine participation of ethnic minorities in clinical research, and (2) proposes strategies that mitigate such barriers. MethodsA systematic literature review explored barriers affecting participation rates from individuals from diverse ethnic backgrounds. An exploratory phase involved two online surveys (researchers and general population) and focus groups (general population) analysed using thematic analysis. ResultsThe systematic review found little published evidence, with most studies undertaken in the USA and focused on specific clinical areas. The exploratory phase showed a discordance between researchers and general publics perspectives on both drivers and barriers to early phase trial participation. DiscussionThese findings were synthesised into a Clinical Trials Participatory Framework, which contextualises reasons for reduced trial participation, while providing mechanisms/strategies to increase uptake among minority ethnic participants. This may guide researchers when implementing strategies to aid under-representation in their samples. Further research should evaluate the framework by actively implementing, testing, and iterating upon the strategies.

8
Evidence Supporting EMA Drug Approvals (2020-2023): A Cross-Sectional Study of Trial Design and Outcomes

Siebert, M.; Caquelin, L.; Naudet, F.; Ross, J. S.; Ramachandran, R.

2026-02-05 health policy 10.64898/2026.02.04.26345500 medRxiv
Top 0.1%
15.8%
Show abstract

BackgroundThe strength and transparency of clinical trial evidence supporting drug approvals has become increasingly scrutinized, particularly considering the increased use of regulatory flexibility and expedited pathways. While U.S. Food and Drug Administration (FDA) standards have been extensively analyzed, evidence standards at the European Medicines Agency (EMA) remain less well-characterized. Thus, this study aims to systematically assess the design, quality, and outcomes of pivotal efficacy trials supporting EMA drug approvals between 2020 and 2023. MethodsWe conducted a cross-sectional analysis of new medicines and biosimilars receiving positive opinions from the EMAs Committee for Medicinal Products for Human Use (CHMP) and subsequent approval by the European Commission between January 2020 and December 2023. Data were extracted from European Public Assessment Reports (EPARs) and EMA medicine databases. Key variables included trial design features, primary endpoint type and achievement status, and justification for approval in cases of failed efficacy endpoints. ResultsBetween 2020 and 2023, 232 drugs were approved by the EMA for 281 indications. Of these, 205 (88.4%) were new active substances and 65 (28.0%) were granted orphan designation. Forty-six products (19.8%) were approved via a special regulatory program, most commonly Conditional Approval (26 products; 11.2%). Cancer was the leading therapeutic area, accounting for 61 approvals (26.3%). Approvals were supported by 393 pivotal clinical trials. Of these, 327 (83.2%) were randomized controlled trials (RCTs) and 218 (66.6% of RCTs) had a superiority design. A total of 232/393 trials (59.0%) relied on surrogate endpoints. Overall, 22 approvals (9.5%) were supported by at least one pivotal trial in which at least one primary endpoint was not met; in seven of these cases (31.8%), the failed trial was the sole pivotal trial. The most common rationale for approval despite null primary results was reliance on the totality of evidence, secondary endpoints, or clinical judgment (9 products; 40.9%). ConclusionsOur findings reveal substantial variability in the design and evidentiary strength of pivotal trials supporting EMA approvals between 2020 and 2023. While the majority of studies were RCTs, reliance on surrogate endpoints was common. That 10% of approvals were based on pivotal trials with null primary endpoints highlights the nuanced role of regulatory judgment in therapeutic evaluation. These findings prompt reflection on evolving evidence standards in drug regulation and underscore the need for transparency and consistent justifications.

9
Adjusting Power Calculations and Trial Sample Size for Treatment Resistance

Tepekule, B.

2025-10-17 health informatics 10.1101/2025.10.14.25337997 medRxiv
Top 0.1%
15.2%
Show abstract

Statistical calculations for clinical trials traditionally assume that if a treatment fails, it fails for mechanistic reasons--the drug itself is ineffective. However, patients may be treatment-resistant, rendering them unable to benefit from an otherwise effective treatment. This creates an identifiability problem: a null hypothesis that we fail to reject can indicate either an ineffective treatment, or an effective treatment tested in a population dominated by treatment-resistant subjects. However, the strategy to administer the drug should be different for these cases. Here we present a simple way to adjust the sample size of a randomized controlled trial to account for the anticipated level of treatment resistance to reach a certain statistical power. We show that the resistant-adjusted population size exponentially increases with the anticipated resistance prevalence, whereas power decreases almost linearly for a given population size as the resistance prevalence increases.

10
From Protocol to Analysis Plan: Development and Validation of a Large Language Model Pipeline for Statistical Analysis Plan Generation using Artificial Intelligence (SAPAI)

Jafari, H.; Chu, P.; Lange, M.; Maher, F.; Glen, C.; Pearson, O. J.; Burges, C.; Martyn, M.; Cross, S.; Carter, B.; Emsley, R.; Forbes, G.

2026-03-19 health systems and quality improvement 10.64898/2026.03.19.26348626 medRxiv
Top 0.1%
14.9%
Show abstract

Background: Statistical Analysis Plans (SAPs) are essential for trial transparency and credibility but are resource-intensive to produce. While Large Language Models (LLMs) have shown promise in drafting protocols, their ability to generate high-quality, protocol-compliant SAPs remains untested against current content guidance. This study developed and validated an LLM-based pipeline for drafting SAPs from clinical trial protocols. Methods: We developed a structured, section-by-section prompting pipeline aligned with standard SAP guidance. We applied this pipeline to nine clinical trial protocols using three leading LLMs: OpenAI GPT-5, Anthropic Claude Sonnet 4, and Google Gemini 2.5 Pro. The resulting 27 SAPs were evaluated against a 46-item quality checklist derived from the published SAP guidelines. Items were double-scored by independent trial statisticians on a 0 to 3 scale for accuracy. We compared performance across LLMs and between item types (descriptive vs. statistical reasoning) using mixed-effects logistic regression. Results: Across 9 trials, the models produced SAP drafts with high overall accuracy (77% to 78%), with no difference in performance between the three LLMs (p=0.79) but varied by content type (p < 0.001). All models performed well on descriptive items (e.g., administrative details, trial design), with lower accuracy for items requiring statistical reasoning (e.g., modelling strategies, sensitivity analyses). Accuracy for statistical items ranged from 67% to 72%, whereas descriptive items achieved 81% to 83% accuracy. Qualitatively, models were prone to specific failure modes in complex sections, such as omitting necessary details for secondary outcome models or hallucinating sensitivity analyses. Discussion: Current LLMs can effectively draft portions of SAPs, offering the potential for substantial time savings in trial documentation. However, a human-in-the-loop approach remains mandatory; while models demonstrate strong capability in producing descriptive content, their independent application to complex statistical methodology design still requires further methodological development and training. Future work should explore advanced prompt engineering, such as retrieval-augmented generation or agentic workflows, to improve reasoning capabilities.

11
A Critical Interpretive Synthesis to Develop Quality Assessment Tools for E-Cigarette Systematic Reviews: Scope and Protocol

O'Leary, R.; Costanzo, F.

2020-05-29 addiction medicine 10.1101/2020.05.25.20112524 medRxiv
Top 0.1%
13.4%
Show abstract

One component of a systematic review is the quality assessment of studies to determine their inclusion or exclusion. Studies on e-cigarettes are conducted in the contentious atmosphere surrounding tobacco harm reduction, which has resulted at times in research bias. Therefore, the quality assessment of studies on e-cigarettes requires more scrutiny than what is provided by generic tools on study design. This topic-specific quality assessment must examine the tests, measurements, and analysis methods used for their adherence to research standards. Furthermore, the studies need to be carefully screened for bias. Because standard quality assessment tools do not provide this topic-specific guidance, we propose to develop quality assessment tools specifically for reviews on e-cigarettes, and for our living systematic reviews on e-cigarettes for tobacco harm reduction.

12
Comparative Effectiveness of Single vs. Dual WhatsApp Reminders on No-shows: A Target Trial Emulation within the Public Health System of Buenos Aires, Argentina.

Esteban, S.; Quintana, G.; Sanchez, M.; Szmulewicz, A.

2026-08-19 health systems and quality improvement 10.64898/2026.08.17.26360609 medRxiv
Top 0.1%
13.1%
Show abstract

Background: Digital reminders reduce outpatient no-shows, but the optimal timing and frequency of messages remain unclear, particularly in Latin American public health systems. We emulated a target trial to evaluate the comparative effectiveness of four WhatsApp reminder strategies on appointment absenteeism and patient-initiated cancellations. Methods: We analyzed administrative and electronic health-record data from the public health system of the Autonomous City of Buenos Aires, Argentina (June 2023-May 2024). Eligible individuals had scheduled an in-person outpatient appointment in one of 15 prioritized specialties at least 75 hours in advance and had a mobile phone on record. We compared four strategies: (1) dual reminders at ~72 and ~24 hours before the appointment; (2) a single reminder at ~72 hours; (3) a single reminder at ~24 hours; and (4) no reminders. The primary outcome was the proportion of no-shows by the end of follow-up. Secondary outcomes were the cumulative incidence of patient-initiated cancellations overall, within 12 hours of the appointment, and followed by rebooking. We emulated the target trial using a cloning-censoring-weighting approach to estimate per-protocol controlled direct effects, with inverse-probability weights to address time-varying confounding and selection bias. Cumulative incidence of secondary outcomes was estimated using weighted Kaplan-Meier curves. Three pre-specified sensitivity analyses and standardized mean differences assessed robustness and covariate balance. Results: A total of 475,214 first eligible person-appointments were included; baseline no-show risk in the control arm was 34.6%. All three active strategies reduced no-shows compared with no reminders. The single 24-hour reminder produced the largest reduction (Risk Ratio [RR] 0.76, 95% CI 0.72, 0.81; Risk Difference [RD] -8.21 percentage points [pp], 95% CI -9.68, -6.54), followed by the dual-reminder strategy (RR 0.80, 95% CI 0.79,0.81; RD -7.05 pp, 95% CI -7.41, -6.71) and the single 72-hour reminder (RR 0.91, 95% CI 0.84,0.99; RD -3.16 pp, 95% CI -5.69, -0.49). All active strategies increased patient-initiated cancellations relative to control, with the dual-reminder strategy producing the largest increase. Sensitivity analyses preserved the qualitative ranking of strategies across all specifications. Conclusions: In this large target trial emulation, a single just-in-time WhatsApp reminder sent ~24 hours before the appointment was as effective as a dual-reminder schedule in preventing no-shows and superior to a distal 72-hour reminder alone. Adding a second, distal reminder provided no measurable benefit for attendance but substantially increased patient-initiated cancellations, which may be operationally valuable when active slot reallocation is a goal. These findings support timing, rather than frequency, as the primary lever of digital-reminder effectiveness, and favor the deployment of a single proximal reminder as the default strategy in resource-constrained outpatient settings.

13
A longitudinal cohort study comparing clinical trials registered on ClinicalTrials.gov that stopped during the first three years of the SARS-CoV-2 pandemic with trials that stopped in the three years prior

Carlisle, B. G.; Hutchinson, N.; Moyer, H.

2026-05-22 public and global health 10.64898/2026.05.20.26353581 medRxiv
Top 0.1%
12.5%
Show abstract

Background: The global SARS-CoV-2 pandemic disrupted healthcare systems worldwide, raising concerns about its impact on clinical research. Early reports suggested reductions in participant enrollment, interruptions to ongoing trials, and challenges to protocol adherence, yet the magnitude and duration of these operational disruptions remain unclear. Methods: We conducted a registry-based analysis comparing clinical trials during the COVID-19 pandemic (December 2019 to November 2022) with a matched pre-pandemic cohort (December 2016 to November 2019). Studies were included if they reported any modifications to trial status, enrollment, or protocols during the study periods. Key variables included trial stoppage, enrollment changes, and adoption of remote or hybrid procedures. Results: The global SARS-CoV-2 pandemic resulted in widespread disruptions to trial operations with 13,323 clinical trials terminated, suspended or withdrawn over the course of the pandemic, a 38% increase compared to the 9,665 trials that stopped in the 3 years prior to the pandemic. Registries indicated a sharp decline in new participant enrollment across geographic regions and therapeutic areas, with partial recovery in later months. Review findings highlighted barriers including patient inaccessibility, staff redeployment, and supply chain interruptions. Conclusions: The pandemic caused system-wide operational shocks that compromised trial timelines and may have downstream methodological consequences. Recovery in enrollment does not imply restoration of pre-pandemic protocol fidelity or outcome ascertainment. Standardized reporting of disruptions, proactive contingency planning, and resilient trial designs are needed to maintain data integrity during large-scale disruptions and to support reliable evidence generation.

14
Was CED the Right Choice? A Decision-Theoretic Evaluation of the CMS Cover with Evidence Development Policy for Aducanumab

Popp, J.; Jutkowitz, E.; Trikalinos, T.

2024-02-14 health policy 10.1101/2024.02.13.24302771 medRxiv
Top 0.1%
11.9%
Show abstract

BackgroundIn 2022, the Centers for Medicare & Medicaid Services (CMS) issued its final national coverage policy for aducanumab, a novel FDA-approved treatment for Alzheimers disease, deciding to Cover with Evidence Development (CED). CMS will thus only pay for the treatment of AD patients enrolled in an approved randomized controlled trial (RCT). We sought to understand whether, given current evidence, CED was best from a societal perspective. MethodsWe conducted a modeling-based expected value of sample information analysis to estimate the expected net decision-theoretic value of a further RCT to evaluate the clinical efficacy of high-dose (10 mg/kg) aducanumab and to determine what sized trial, if any, is optimal conditional on an initial decision to cover or not. We also evaluated the expected net benefit of the manufacturers proposed RCT ( ENVISION). We considered two post-trial decision criteria: cost-effectiveness given updated evidence ( efficiency) and does the new trial demonstrate a statistical significant (p<0.05) clinical benefit. Results were used to calculate the expected population net monetary benefit (NMB) of four decision alternatives (including CED) depending on an initial coverage and trial decision. We ranked alternatives and calculated the expected opportunity loss of a suboptimal decision. We used a societal perspective and focused on willingness-to-pay (WTP) values for a quality-adjusted life year (QALY) between $50K-$200K. We conducted scenario analyses using different assumptions about population size, efficacy, and drug cost. FindingsCMSs decision to not cover aducanumab avoids an expected societal loss (NMB) of $15B-$110B. Even an optimally designed RCT would confer no or negative decision-theoretic value for WTP[&le;]$100K or with statistical significance as a post-trial decision criterion, respectively, and thus denying coverage without a trial (rather than CED) is clearly preferable. For WTP=$150K (WTP=$200K) and assuming an efficiency criterion, CED with ENVISION or a similar trial is reasonable (decidedly optimal). The case for future research would become less ambiguous if the manufacturer again voluntarily dropped the price [&ge;]50%. InterpretationThe societal net value of a future trial (and thus CED) depends on how CMS would use the trial results to update its coverage decision and the WTP per QALY. Assuming CMS policymakers can avoid the pitfalls of a legal framework that limits their ability to consider costs in coverage decisions, the CED decision is at least reasonable, if not optimal, if a QALY is valued [&ge;]$150K.

15
Evaluation of Clinical Trial Data Sharing Policy in Leading Medical Journals

Danchev, V.; Min, Y.; Borghi, J.; Baiocchi, M.; Ioannidis, J. P. A.

2020-05-11 health policy 10.1101/2020.05.07.20094656 medRxiv
Top 0.1%
11.9%
Show abstract

BackgroundThe benefits from responsible sharing of individual-participant data (IPD) from clinical studies are well recognized, but stakeholders often disagree on how to align those benefits with privacy risks, costs, and incentives for clinical trialists and sponsors. Recently, the International Committee of Medical Journal Editors (ICMJE) required a data sharing statement (DSS) from submissions reporting clinical trials effective July 1, 2018. We set out to evaluate the implementation of the policy in three leading medical journals (JAMA, Lancet, and New England Journal of Medicine (NEJM)). MethodsA MEDLINE/PubMed search of clinical trials published in the three journals between July 1, 2018 and April 4, 2020 identified 487 eligible trials (JAMA n = 112, Lancet n = 147, NEJM n = 228). Two reviewers evaluated each of the 487 articles independently. Captured outcomes were declared data availability, data type, access, conditions and reasons for data (un)availability, and funding sources. Findings334 (68.6%, 95% confidence interval (CI), 64.1%-72.5%) articles declared data sharing, with non-industry NIH-funded trials exhibiting the highest rates of declared data sharing (88.9%, 95% CI, 80.0%-97.8) and industry-funded trials the lowest (61.3%, 95% CI, 54.3%-68.3). However, only two IPD datasets were actually deidentified and publicly available as of April 10, 2020. The remaining were supposedly accessible via request to authors (42.8%, 143/334), repository (26.6%, 89/334), and company (23.4%, 78/334). Among the 89 articles declaring to store IPD in repositories, only 17 articles (19.1%) deposited data, mostly due to embargo and regulatory approval. Embargo was set in 47.3% (158/334) of data-sharing articles, and in half of them the period exceeded 1 year or was unspecified. InterpretationMost trials published in JAMA, Lancet, and NEJM after the implementation of the ICMJE policy declared their intent to make clinical data available. However, a wide gap between declared and actual data sharing exists. To improve transparency and data reuse, journals should promote the use of unique pointers to dataset location and standardized choices for embargo periods and access requirements. All data, code, and materials used in this analysis are available on OSF at https://osf.io/s5vbg/.

16
Clinical Trial Data Sharing: A Cross-Sectional Study of Outcomes Associated with Two NIH Models

Rowhani-Farid, A.; Egilman, A. C.; Zhang, A. D.; Gross, C. P.; Krumholz, H. M.; Ross, J. S.

2021-09-13 health policy 10.1101/2021.09.10.21263404 medRxiv
Top 0.1%
11.9%
Show abstract

BackgroundThe impact and value of clinical trial data sharing, including the number and quality of publications that result from shared data - "shared data publications" - may differ depending on the data sharing model used. MethodsWe characterized the outcomes associated with two data sharing models previously used by Institutes of the U.S. National Institutes of Health (NIH): NHLBIs centralized model, which uses a repository to manage data sharing requests, and NCIs decentralized model, which entrusted research groups to independently manage data sharing requests. We identified trials completed in 2010 that met NIH data sharing criteria and matched studies sponsored by each Institute based on cost or size, determining whether trial data were shared and the frequency of shared data publications. ResultsWe identified 14 NHLBI-funded trials and 48 NCI-funded trials that met NIH data sharing criteria. We matched 14 NCI-funded trials to the 14 NHLBI-funded trials; among these, 4 NHLBI-sponsored trials (29%) and 2 NCI-sponsored trials (14%) shared data. From the 2 NCI-sponsored trials sharing data, we identified 2 shared data publications, one per trial, both of which were meta-analyses. From the 4 NHLBI-sponsored trials sharing data, we identified 7 shared data publications, all using data from 1 trial, 5 of which were pooled analyses and 2 reported secondary outcomes. ConclusionWhen characterizing the outcomes associated with two NIH data sharing models, both the NHLBI and the NCI models resulted in only 21% of trials sharing data and few shared data publications. There are opportunities to optimize clinical trial data sharing efforts both to enhance clinical trial data sharing and increase the number of shared data publications.

17
AI-Assisted Data Extraction with a Large Language Model: A Study Within Reviews

Gartlehner, G.; Kugley, S.; Crotty, K.; Viswanathan, M.; Dobrescu, A.; Nussbaumer-Streit, B.; Booth, G.; Treadwell, J.; Han, J. M.; Wagner, J.; Apaydin, E.; Coppola, E.; Maglione, M.; Hilscher, R.; Chew, R.; Pilar, M.; Swanton, B.; Kahwati, L.

2025-03-21 health systems and quality improvement 10.1101/2025.03.20.25324350 medRxiv
Top 0.1%
11.8%
Show abstract

BackgroundData extraction is a critical but error-prone and labor-intensive task in evidence synthesis. Unlike other artificial intelligence (AI) technologies, large language models (LLMs) do not require labeled training data for data extraction. ObjectiveTo compare an AI-assisted to a traditional y data extraction process. DesignStudy within reviews (SWAR) utilizing a prospective, parallel group comparison with blinded data adjudicators. SettingWorkflow validation within six ongoing systematic reviews of interventions under real-world conditions. InterventionInitial data extraction using an LLM (Claude versions 2.1, 3.0 Opus, and 3.5 Sonnet) verified by a human reviewer. MeasurementsConcordance, time on task, accuracy, recall, precision, and error analysis. ResultsThe six systematic reviews of the SWAR contributed 9,341 data elements, extracted from 63 studies. Concordance between the two methods was 77.2%. The accuracy of the AI-assisted approach compared with enhanced human data extraction was 91.0%, with a recall of 89.4% and a precision of 98.9%. The AI-assisted approach had fewer incorrect extractions (9.0% vs. 11.0%) and similar risks of major errors (2.5% vs. 2.7%) compared to the traditional human-only method, with a median time saving of 41 minutes per study. Missed data items were the most frequent errors in both approaches. LimitationsAssessing the concordance of data extractions and classifying errors required subjective judgment. Tracking time on task consistently was challenging. ConclusionThe use of an LLM can improve accuracy of data extraction and save time in evidence synthesis. Results reinforce previous findings that human-only data extraction is prone to errors. Primary Funding SourceUS Agency for Healthcare Research and Quality, RTI International RegistrationSWAR28 Gerald Gartlehner (2023 FEB 11 2102).pdf

18
Statistical Analysis Plan for the Primary and Selected Secondary Endpoints in the ENHANCED-SPS Study

Balzer, L. B.; Kabami, J.; ENHANCED-SPS Study Team,

2023-09-04 hiv aids 10.1101/2023.09.01.23294899 medRxiv
Top 0.1%
11.7%
Show abstract

This document provides the statistical analytic plan (SAP) for the ENHANCED-SPS study, a cluster randomized trial to evaluate the effects of peer-led, multicomponent intervention on viral suppression and other care outcomes among pregnant and breastfeeding women with HIV in South Western Uganda (Clinicaltrials.gov: NCT04122144). The SAP was locked prior to unblinding and effect estimation. This SAP was embargoed until August 31, 2023 when it was submitted to medRxiv.

19
Two Stage Designs for Phase III Trials

Follmann, D.; Proschan, M.

2020-08-03 epidemiology 10.1101/2020.07.29.20164525 medRxiv
Top 0.1%
11.5%
Show abstract

Phase III platform trials are increasingly used to evaluate a sequence of treatments for a specific disease. Traditional approaches to structure such trials tend to focus on the sequential questions rather than the performance of the entire enterprise. We consider two-stage trials where an early evaluation is used to determine whether to continue with an individual study. To evaluate performance, we use the ratio of expected wins (RW), that is, the expected number of reported efficacious treatments using a two-stage approach compared to that using standard phase III trials. We approximate the test statistics during the course of a single trial using Brownian Motion and determine the optimal stage 1 time and type I error rate to maximize RW for fixed power. At times, a surrogate or intermediate endpoint may provide a quicker read on potential efficacy than use of the primary endpoint at stage 1. We generalize our approach to the surrogate endpoint setting and show improved performance, provided a good quality and powerful surrogate is available. We apply our methods to the design of a platform trial to evaluate treatments for COVID-19 disease.

20
The Dearth of Representation in FDA Approved Drug Trials

Ibilibor, C.; Armbruster, S.; Parker, R.; Yu, J.-R.; Barros, A.

2024-01-17 health policy 10.1101/2024.01.16.24301376 medRxiv
Top 0.1%
11.0%
Show abstract

The generalizability of data derived from randomized controlled trials is of paramount importance given their utility in the Food & Drug Administration (FDA) drug approval process. An essential part of this process is the inclusion of reliably reported gender, race and ethnicity data in trials that lead to FDA drug approval. Despite previous mandates by the FDA and Clinicaltrials.gov, gender and race-specific data remains under reported. We reviewed 100 most recently approved FDA medications, and abstracted the clinical trial data from Clinicaltrials.gov that supported their approval. We then compared these FDA approved trials to non-FDA approved trials from the same year and of similar size. We found that 40% of the FDA trials were missing race/ethnicity information, while 24% of these trials did not include gender information. We demonstrate that there remains a significant amount of missing gender and racial/ethnic data in trials that lead to FDA-approved medications.