Back

Epidemiology

Ovid Technologies (Wolters Kluwer Health)

All preprints, ranked by how well they match Epidemiology's content profile, based on 32 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
What can be learned from viral co-detection studies in human populations

Chin, T.; Foxman, E. F.; Watkins, T. A.; Lipsitch, M.

2023-06-18 epidemiology 10.1101/2023.06.17.23291541 medRxiv
Top 0.1%
52.9%
Show abstract

When respiratory viruses co-circulate in a population, individuals may be infected with multiple pathogens and experience possible virus-virus interactions, where concurrent or recent prior infection with one virus affects the infection process of another virus. While experimental studies have provided convincing evidence for within-host mechanisms of virus-virus interactions, evaluating evidence for viral interference or potentiation using population-level data has proven more difficult. Recent studies have quantified the prevalence of co-detections using populations drawn from clinical settings. Here, we focus on selection bias issues associated with this study design. We provide a quantitative account of the conditions under which selection bias arises in these studies, review previous attempts to address this bias, and propose unbiased study designs with sample size estimates needed to ascertain viral interference. We show that selection bias is expected in cross-sectional co-detection prevalence studies conducted in clinical settings, except under a strict set of assumptions regarding the relative probabilities of having symptoms under different viral states. Population-wide studies that sample participants irrespective of their symptom status would meanwhile require large sample sizes to be sufficiently powered to detect viral interference, suggesting that a studys timing, inclusion criteria, and the expected magnitude of interference are instrumental in determining feasibility.

2
Vaccine efficacy against naturally asymptomatic infections: A novel estimand for quantifying vaccine effects

Rogawski McQuade, E. T.; Nabi, R.; Codi, A.; Dean, N.; Lipsitch, M.; Benkeser, D.

2025-08-24 epidemiology 10.1101/2025.08.20.25334070 medRxiv
Top 0.1%
39.0%
Show abstract

The naive approach to estimating the effects of a vaccine on asymptomatic infections, which compares the risk of asymptomatic infection among vaccinated and unvaccinated individuals, can be misleading because it is comprised of two effects: the vaccine preventing asymptomatic infections and the vaccine converting symptomatic to asymptomatic infections. When the latter effect is strong, vaccines can appear harmful with respect to asymptomatic infections. Using a causal principal stratification framework, we formalize an estimand, vaccine efficacy against naturally asymptomatic infection (VENAI), that describes the effectiveness of a vaccine in preventing asymptomatic infections among individuals who would naturally (i.e., in the absence of vaccine) be expected to be asymptomatic. This estimand excludes vaccine effects that convert symptomatic cases to asymptomatic infections, and we demonstrate how this makes it a more natural analogue of the usual vaccine efficacy estimands against infection and symptomatic disease. We describe the assumptions under which this estimand can be identified and estimated from randomized and observational studies. We further identify and estimate bounds that do not require cross-world independence assumptions and characterize sensitivity analyses around the main assumption needed for identification. Finally, we apply these methods to a randomized trial of the COVID-19 mRNA-1273 vaccine. In this trial, VENAI was higher than standard estimates of efficacy against asymptomatic infections and was similar in magnitude to efficacy against any infection. Reporting VENAI in vaccine trials in addition to other vaccine effects would improve interpretability, could broaden understanding of vaccine impact on transmission, and provide insights into immunological mechanisms.

3
Adjusting for hidden biases in sexual behaviour data: a mechanistic approach

Knight, J.; Wang, S.; Mishra, S.

2023-08-20 epidemiology 10.1101/2023.08.16.23294164 medRxiv
Top 0.1%
31.5%
Show abstract

BackgroundTwo required inputs to mathematical models of sexually transmitted infections are the average duration in epidemiological risk states (e.g., selling sex) and the average rates of sexual partnership change. These variables are often only available as aggregate estimates from published cross-sectional studies, and may be subject to distributional, sampling, censoring, and measurement biases. MethodsWe explore adjustments for these biases using aggregate estimates of duration in sex work and numbers of reported sexual partners from a published 2011 survey of female sex worker in Eswatini. We develop adjustments from first principles, and construct Bayesian hierarchical models to reflect our mechanistic assumptions about the bias-generating processes. ResultsWe show that different mechanisms of bias for duration in sex work may "cancel out" by acting in opposite directions, but that failure to consider some mechanisms could over- or underestimate duration in sex work by factors approaching 2. We also show that conventional interpretations of sexual partner numbers are biased due to implicit assumptions about partnership duration, but that unbiased estimators of partnership change rate can be defined that explicitly incorporate a given partnership duration. We highlight how the unbiased estimator is most important when the survey recall period and partnership duration are similar in length. ConclusionsWhile we explore these bias adjustments using a particular dataset, and in the context of deriving inputs for mathematical modelling, we expect that our approach and insights would be applicable to other datasets and motivations for quantifying sexual behaviour data.

4
Some principles for using epidemiologic study results to parameterize transmission models

Joshi, K.; Kahn, R.; Boyer, C.; Lipsitch, M.

2023-10-03 infectious diseases 10.1101/2023.10.03.23296455 medRxiv
Top 0.1%
31.3%
Show abstract

BackgroundInfectious disease models, including individual based models (IBMs), can be used to inform public health response. For these models to be effective, accurate estimates of key parameters describing the natural history of infection and disease are needed. However, obtaining these parameter estimates from epidemiological studies is not always straightforward. We aim to 1) outline challenges to parameter estimation that arise due to common biases found in epidemiologic studies and 2) describe the conditions under which careful consideration in the design and analysis of the study could allow us to obtain a causal estimate of the parameter of interest. In this discussion we do not focus on issues of generalizability and transportability. MethodsUsing examples from the COVID-19 pandemic, we first identify different ways of parameterizing IBMs and describe ideal study designs to estimate these parameters. Given real-world limitations, we describe challenges in parameter estimation due to confounding and conditioning on a post-exposure observation. We then describe ideal study designs that can lead to unbiased parameter estimates. We finally discuss additional challenges in estimating progression probabilities and the consequences of these challenges. ResultsCausal estimation can only occur if we are able to accurately measure and control for all confounding variables that create non-causal associations between the exposure and outcome of interest, which is sometimes challenging given the nature of the variables we need to measure. In the absence of perfect control, non-causal parameter estimates should still be used, as sometimes they are the best available information we have. ConclusionsIdentifying which estimates from epidemiologic studies correspond to the quantities needed to parameterize disease models, and determining whether these parameters have causal interpretations, can inform future study designs and improve inferences from infectious disease models. Understanding the way in which biases can arise in parameter estimation can inform sensitivity analyses or help with interpretation of results if the magnitude and direction of the bias is understood.

5
Use of the test-negative design to estimate the protective effect of a scalar immune measure: A simulation analysis

Zhang, Z.; Boyer, C.; Lipsitch, M.

2024-11-23 epidemiology 10.1101/2024.11.22.24317757 medRxiv
Top 0.1%
31.0%
Show abstract

BackgroundThe relationship between antibody levels (more generally, a scalar measure of immune protection) at the time of exposure to infection (so-called exposure-proximal correlates of protection) and the risk of infection given exposure is of central interest in evaluating the evolution of immune protection conferred by prior infection and/or vaccination. A version of the test-negative study design (TND), adapted from vaccine effectiveness studies, has been used to assess this relationship. However, the conditions under which such a study identifies the relationship between immune measurements and protection have not been defined. ObjectiveTo evaluate the conditions for TNDs to estimate the relationship between antibody levels or a similar scalar measurement of immunity (hereafter exposure-proximal correlates of protection, COP) and the relative incidence rate of infection given exposure. MethodIndividual-based transmission models, linking infection risk linearly and nonlinearly with COP value and accounting for waning immunity post-vaccination and -infection, were used. Simulations were performed of a TND with sampling on predetermined dates. Data from either one or multiple simulation days were analyzed using logistic regression and generalized additive models. ResultA correctly specified logistic regression model provided an unbiased estimate of the effectiveness of specific COP levels (analogous to vaccine effectiveness). Aggregating data across different simulation dates with incidence-density sampling also provided reliable estimates of protection. When, as is generally the case, the functional form relating COP level to protection is unknown, generalized additive models offer a more flexible alternative to traditional logistic regression approaches. ConclusionA TND can validly estimate the relative effect of an immune COP at the time of exposure on the incidence rate of infection via logistic regression if the functional form of the effect is known and appropriately modeled or unknown a semiparametric approach. Future research should further examine the dynamics of immunity waning and boosting for more reliable inference.

6
Assessment of the effectiveness of required weekly COVID-19 surveillance antigen testing at a university

Ryan, C. W.

2023-11-23 epidemiology 10.1101/2023.11.22.23298917 medRxiv
Top 0.1%
30.7%
Show abstract

ObjectivesTo mitigate the COVID-19 pandemic, many institutions implemented a regimen of periodic required testing, irrespective of symptoms. The effectiveness of this"surveillance testing" requires assessment. MethodsI fit a zero-inflated negative binomial model to COVID-19 testing and case investigation data between 1 November 2020 and 15 May 2021, from young adult subjects in one community. I compared the duration of symptoms at time of specimen collection in those diagnosed via (1) surveillance testing at a university, (2) the same universitys student health services, and (3) all other testing venues. ResultsThe data comprised 2926 records: 393 from surveillance testing, 493 from student health service, and 2040 from other venues. About 65% of people with COVID-19 detected via surveillance testing were already symptomatic at time of specimen collection. Predicted mean duration of pre-testing symptoms was 1.7 days (95% CI 1.59 to 1.84) for the community, 1.81 days (95% CI 1.62 to 1.99) for surveillance, and 2 days (95% CI 1.83 to 2.16) for student health service. The modelled "inflated" proportions of asymptomatic subjects from the surveillance stream and the other/community stream were comparable (odds ratio 0.95, p = 0.7709). Comparing surveillance testing with the student health service, the proportion of "excess" zero symptom durations was signficantly higher in the former (Chi-square = 12.08, p = 0.0005) ConclusionsSurveillance testing at a university detected 393 people with COVID-19, but no earlier in their trajectory than similar-aged people detected in the broader community. This casts some doubt on the public health value of such programs, which tend to be labor-intensive and expensive. 2 Three-question summary boxO_ST_ABSWhat is the current understanding of this subject?C_ST_ABSAssessments of long-term operational effectiveness of COVID-19 "surveillance testing" have not been published. What does this report add to the literature?During the 2020-2021 academic year at one university, people with COVID-19 detected via compulsory weekly surveillance antigen testing were equally likely to be symptomatic at time of detection, and for just as long, as similar-aged people detected via testing venues in the community. What are the implications for public health practice?Surveillance testing programs during the pandemic consumed a large amount of time, money, and effort. In future respiratory pandemics, resources might be better devoted to other mitigation measures.

7
Causal Inference via Electronic Health Records in the National Clinical Cohort Collaborative: Challenges and Solutions in Long COVID Research

Butzin-Dozier, Z.; Ji, Y.; Wang, L.-C.; Anzalone, A. J.; Hurwitz, E.; Patel, R. C.; van der Laan, M.; Colford, J. M.; Hubbard, A. E.; on behalf of the N3C Consortium,

2025-06-08 epidemiology 10.1101/2025.06.06.25329168 medRxiv
Top 0.1%
28.7%
Show abstract

Observational analyses of electronic health record (EHR) data using databases such as the National Clinical Cohort Collaborative include unique challenges for researchers seeking causal inferences, particularly when evaluating subjectively-defined outcomes like Long COVID. We explore several challenges and describe potential solutions. 1. Lack of true negatives: Many diagnoses and conditions either have a positive indicator or a missing status, requiring investigators to carefully consider which patients are likely negative for this condition. 2. Differential monitoring: EHR data include nonrandom missingness driven by patients engaging with the healthcare system at different rates, which is often related to both the exposure and outcome of interest. 3. Bias: EHR data sources face many biases, but are particularly vulnerable to informative missingness, differential monitoring, and model misspecification. 4. Large sample size: High precision (i.e., narrow confidence intervals) paired with potential bias leads to a high risk of incorrectly rejecting the null hypothesis. 5. Defining index time: It is important that investigators deliberately define index time (i.e., t0, baseline) to ensure that they only adjust for baseline confounders and do not adjust for (or condition on) factors that are affected by the exposure of interest (i.e., colliders or mediators). 6. Parameter selection: Investigators should only select parameters that are supported by the data distribution. This manuscript provides an overview of these challenges and solutions, using both simulated data and real-world data, with the outcome of Long COVID as the running example.

8
Accounting for Twins and Other Multiple Births in Perinatal Studies Conducted Using Healthcare Administration Data

Brown, J. P.; Yland, J. J.; Williams, P. L.; Huybrechts, K. F.; Hernandez-Diaz, S.

2024-01-24 epidemiology 10.1101/2024.01.23.24301685 medRxiv
Top 0.1%
26.7%
Show abstract

The analysis of perinatal studies is complicated by twins and other multiple births even when they are not the exposure, outcome, or a confounder of interest. Common approaches to handling multiples in studies of infant outcomes include restriction to singletons, counting outcomes at the pregnancy-level (i.e., by counting if at least one twin experienced a binary outcome), or infant-level analysis including all infants and, typically, accounting for clustering of outcomes by using generalised estimating equations or mixed effects models. Several healthcare administration databases only support restriction to singletons or pregnancy-level approaches. For example, in MarketScan insurance claims data, diagnoses in twins are often assigned to a single infant identifier, thereby preventing ascertainment of infant-level outcomes among multiples. Different approaches correspond to different causal questions, produce different estimands, and often rely on different assumptions. We demonstrate the differences that can arise from these different approaches using Monte Carlo simulations, algebraic formulas, and an applied example. Furthermore, we provide guidance on the handling of multiples in perinatal studies when using healthcare administration data.

9
Defining and emulating target trials of the effects of postexposure vaccination using observational data

Lipsitch, M.; Boyer, C.

2023-05-05 epidemiology 10.1101/2023.05.03.23289471 medRxiv
Top 0.1%
26.4%
Show abstract

Postexposure vaccination has the potential to prevent or modify the course of clinical disease among those exposed to a pathogen. However, due to logistical constraints, postexposure vaccine trials have been difficult to implement in practice. In place of trials, investigators have used observational data to estimate the effectiveness or optimal timing window for postexposure vaccines, but the relationship between these analyses and those that would be conducted in a trial is often unclear. Here, we define several possible target trials for postexposure vaccination and show how, under certain conditions, they can be emulated using observational data. We emphasize the importance of the incubation period and the timing of vaccination in trial design and emulation. As an example, we specify a protocol for postexposure vaccination against mpox and provide a step-by-step description of how to emulate it using data from a healthcare database or contact tracing program. We further illustrate some of the benefits of the target trial approach through simulation.

10
Design and Estimation for the Population Prevalence of Infectious Diseases

Oh, E. J.; Mikytuck, A.; Lancaster, V.; Goldstein, J.; Keller, S.

2021-02-08 epidemiology 10.1101/2021.02.05.21251231 medRxiv
Top 0.1%
22.9%
Show abstract

Understanding the prevalence of infections in the population of interest is critical for making data-driven public health responses to infectious disease outbreaks. Accurate prevalence estimates, however, can be difficult to calculate due to a combination of low population prevalence, imperfect diagnostic tests, and limited testing resources. In addition, strategies based on convenience samples that target only symptomatic or high-risk individuals will yield biased estimates of the population prevalence. We present Bayesian multilevel regression and poststratification models that incorporate probability sampling designs, the sensitivity and specificity of a diagnostic test, and specimen pooling to obtain unbiased prevalence estimates. These models easily incorporate all available prior information and can yield reasonable inferences even with very low base rates and limited testing resources. We examine the performance of these models with an extensive numerical study that varies the sampling design, sample size, true prevalence, and pool size. We also demonstrate the relative robustness of the models to key prior distribution assumptions via sensitivity analyses.

11
Accounting for comorbidity in etiological research

Khachadourian, V.; Janecka, M.

2025-01-21 epidemiology 10.1101/2025.01.19.25320775 medRxiv
Top 0.1%
22.7%
Show abstract

IntroductionDespite the theoretical advancements and recommendations regarding covariate adjustment in causal inference, clinical studies often fail to explicitly state the underlying assumptions related to causal structure among the study variables. Specifically, despite the pervasive nature of comorbidity, explicit causal assumptions about the role of comorbidity in exposure-outcome relationships are often lacking, potentially leading to inappropriate accounting for comorbid conditions and resulting in biased effect estimates. This study aims to explore common causal structures involving comorbidity and provide guidance for handling it in etiologic research. MethodsWe use Directed Acyclic Graphs (DAGs) to depict six causal scenarios involving comorbidity as a confounder, mediator, collider, or consequence of the exposure or outcome. Simulations were conducted across 5,000 iterations for each scenario, assessing the impact of conditioning on comorbidity under three effect measures (mean difference, odds ratio, risk ratio). Bias was evaluated by comparing adjusted and unadjusted effect estimates to the true values. ResultsThe impact of conditioning on comorbidity varied by its causal role. Adjusting for comorbidity mitigated bias when it acted as a confounder, but introduced bias when it was a mediator or collider. In instances where comorbidity was a consequence of either the exposure or outcome, the decision to adjust depended on the research objectives. Nonlinear models revealed differences in marginal and conditional effects due to non-collapsibility. DiscussionExplicit causal assumptions are essential for selecting appropriate analytical strategies in etiologic research. This study provides practical guidance on handling comorbidity-related challenges, highlighting the need for study design and analysis to align with research objectives. Future work should address more complex causal structures and other methodological challenges.

12
Enhanced screening and bacterial sexually-transmitted infection diagnoses after HIV pre-exposure prophylaxis initiation

Parker, A. M.; Jenness, S. M.; Singer, B. J.; Chang, J. J.; Bruxvoort, K. J.; Lewnard, J. A.

2025-12-29 epidemiology 10.64898/2025.12.19.25342713 medRxiv
Top 0.1%
21.9%
Show abstract

BackgroundRecipients of HIV pre-exposure prophylaxis (PrEP) experience higher rates of gonorrhea and chlamydia diagnoses than non-recipients. However, it is unclear if these observations reflect a causal relationship between PrEP initiation and acquisition of sexually transmitted infections (i.e., behavioral "risk compensation"), or alternatively a diagnostic bias related to PrEP recipients screening more frequently than non-recipients. MethodsWe conducted a self-controlled case series study comparing rates of gonorrhea and chlamydia diagnoses after versus before PrEP initiation among commercially-insured US males in the Merative MarketScan(R) Research Databases (2016-2019). We compared matched incidence rate ratio (IRR) estimates after versus before PrEP initiation to model-based expectations under a null hypothesis of no change in infection risk. We derived null expectations via a mathematical model reflecting changes in screening after PrEP initiation, estimating parameters under a Bayesian framework. ResultsIndividuals experienced increased rates of gonorrhea (IRR=3.47 [95% confidence interval: 2.57-4.72]) and chlamydia (IRR=3.59 [2.50-5.27]) diagnoses after PrEP initiation, with IRR estimates differing by anatomical site (IRR for either infection=1.23 [0.71-2.13] and 5.98 [3.63-10.29] for urogenital and extragenital diagnoses, respectively). After accounting for increased rates of screening after PrEP initiation, rates of gonorrhea and chlamydia diagnoses only modestly exceeded expectations under the null hypothesis of no change in infection risk (8-14% greater-than-expected rates of gonorrhea and chlamydia diagnoses, respectively; one-sided p>0.1 for all tests). Overall, we estimated that 80-81% of observed increases in rates of gonorrhea and chlamydia diagnoses after PrEP initiation, respectively, were attributable to increases in asymptomatic screening. ConclusionsOur findings suggest higher-frequency asymptomatic screening, rather than behavioral risk compensation, is the primary driver of increased rates of gonorrhea and chlamydia diagnoses after PrEP initiation.

13
Parallel Trends in an Unparalleled Pandemic: Difference-in-differences for infectious disease policy evaluation

Feng, S.; Bilinski, A.

2024-04-10 infectious diseases Community evaluation 10.1101/2024.04.08.24305335 medRxiv
Top 0.1%
21.1%
Show abstract

Researchers frequently employ difference-in-differences (DiD) to study the impact of public health interventions on infectious disease outcomes. DiD assumes that treatment and non-experimental comparison groups would have moved in parallel in expectation, absent the intervention ("parallel trends assumption"). However, the plausibility of parallel trends assumption in the context of infectious disease transmission is not well-understood. Our work bridges this gap by formalizing epidemiological assumptions required for common DiD specifications, positing an underlying Susceptible-Infectious-Recovered (SIR) data-generating process. We demonstrate that popular specifications can encode strict epidemiological assumptions. For example, DiD modeling incident case numbers or rates as outcomes will produce biased treatment effect estimates unless untreated potential outcomes for treatment and comparison groups come from a data-generating process with the same initial infection and equal transmission rates at each time step. Applying a log transformation or modeling log growth allows for different initial infection rates under an "infinite susceptible population" assumption, but invokes conditions on transmission parameters. We then propose alternative DiD specifications based on epidemiological parameters - the effective reproduction number and the effective contact rate - that are both more robust to differences between treatment and comparison groups and can be extended to complex transmission dynamics. With minimal power difference incidence and log incidence models, we recommend a default of the more robust log specification. Our alternative specifications have lower power than incidence or log incidence models, but have higher power than log growth models. We illustrate implications of our work by re-analyzing published studies of COVID-19 mask policies. Significance StatementDifference-in-differences is a popular observational study design for policy evaluation. However, it may not perform well when modeling infectious disease outcomes. Although many COVID-19 DiD studies in the medical literature have used incident case numbers or rates as the outcome variable, we demonstrate that this and other common model specifications may encode strict epidemiological assumptions as a result of non-linear infectious disease transmission. We unpack the assumptions embedded in popular DiD specifications assuming a Susceptible-Infected-Recovered data-generating process and propose more robust alternatives, modeling the effective reproduction number and effective contact rate.

14
Cluster-weighted modified Poisson regression for estimating risk ratios in longitudinal data with informative cluster sizes

Bather, J. R.; Anyaso-Samuel, S.; Chen, Y.; Elliott, L.; Bennett, A. S.; Goodman, M. S.

2025-05-25 epidemiology 10.1101/2025.05.23.25328253 medRxiv
Top 0.1%
19.8%
Show abstract

Variation in binary outcomes over time by cluster size arises across various biomedical disciplines, including reproductive health, dental medicine, and psychiatric epidemiology. This study formally integrates modified Poisson regression with cluster-weighted generalized estimating equations (MP-CWGEE) for computing risk ratios in longitudinal studies with informative cluster sizes. Using a comprehensive Monte-Carlo simulation study, we empirically evaluated MP-CWGEEs statistical properties against alternative modeling approaches: MP-GEE, log-binomial CWGEE (LB-CWGEE), and log-binomial GEE (LB-GEE). We conducted 1,000 simulations across varying sample sizes, risk ratios, and informativeness degrees. MP-CWGEE demonstrated superior performance in model convergence, empirical bias, average estimated standard error, coverage, and Type 1 error control. While LB-CWGEE showed comparable results, its convergence rates were slightly inferior. The benefits of cluster-weighted models (MP-CWGEE and LB-CWGEE) over unweighted models (MP-GEE and LB-GEE) were pronounced in scenarios with informative cluster sizes. We demonstrated MP-CWGEEs practical application to a cohort study of people who used illicit opioids in New York City. We also provided implementation code for R, Stata, and SAS to facilitate wider adoption of the MP-CWGEE approach.

15
An E-value-Informed Sensitivity Analysis Framework for Hybrid Controlled Trials

Liu, C.; Mayer, M.; Lactaoen, K.; Gomez, L.; Weissman, G.; Hubbard, R.

2026-03-06 epidemiology 10.64898/2026.03.05.26347653 medRxiv
Top 0.1%
19.1%
Show abstract

Hybrid controlled trials (HCTs) incorporate real-world data into randomized controlled trials (RCTs) by augmenting the internal control arm with patients receiving the same treatment in routine care. Beyond increasing power, HCTs may improve recruitment by supporting unequal randomization ratios that increase patient access to experimental treatments. However, HCT validity is threatened by bias from unmeasured confounding due to lack of randomization of external controls, leading to outcome non-exchangeability between internal and external control patients. To address this challenge, we developed a sensitivity analysis framework to assess the robustness of HCT results to potential unmeasured confounding. We propose a tipping point analysis that adapts the E-value framework to the HCT setting where trial participation rather than treatment assignment is subject to confounding. To aid interpretation, we also introduce a data-driven benchmark representing the strength of unmeasured confounding reflected by the observed outcome non-exchangeability. We then propose an operational decision rule and evaluate its performance through simulation studies. Finally, we illustrate the approach using an asthma trial augmented by data from electronic health records. Simulation results demonstrate that our decision rule safeguards against Type I error inflation while preserving the power gains achieved by incorporating external data. In settings where moderate unmeasured confounding led to poorer outcomes for external controls, Type I error was controlled near the nominal 5% level, and power increased by 10-20% compared with analyses using RCT data alone. Our approach provides a practical, interpretable method to assess HCT robustness, supporting rigorous inference when integrating external real-world data.

16
Epidemiologic Moderators of the Effectiveness of Routine Screening for LAIs in High-Biosafety Environments

Cohen, B.; Hanage, W.; Menzies, N. A.; Croke, K.

2026-04-06 epidemiology 10.64898/2026.04.05.26350204 medRxiv
Top 0.1%
18.8%
Show abstract

Justification: Accidental lab-acquired infections (LAIs) with potential pandemic pathogens (PPPs) in high-biosafety research facilities risk causing a pandemic. Routine testing of lab workers for LAIs coupled with isolation of infected workers could reduce the risk, but the impact of such an intervention may depend on pathogens' epidemiological characteristics. Objective: This study aims to understand how the epidemiological characteristics of PPPs moderate the efficacy of a routine testing and isolation intervention in preventing larger outbreaks after an LAI. Methods: We employed a discrete-time stochastic network infectious disease model to run 625,000 epidemic simulations encompassing 625 unique combinations of five parameters of interest: test frequency, pathogen transmissibility, the self-isolation rate for symptomatic cases, the percentage of cases that are asymptomatic, and the percentage of infectious time that is spent in the pre-symptomatic state among those who show symptoms. To summarize the Monte Carlo simulations, we paired visual analysis with logistic regression for formal hypothesis testing, with an emphasis on the interaction terms that capture the moderating effect of epidemiological parameters on the impact of test frequency. Main Results: There were four main findings. First, the relative reductions in risk of outbreak that were caused by increased test frequency were inversely correlated with pathogen transmissibility. Second, the effect of test frequency was magnified at higher asymptomatic shares when the symptomatic self-isolation rate was high, but minimally when the self-isolation rate is low. Third, the direction of how the symptomatic self-isolation rate moderated the effect of increased test frequency depended on the asymptomatic share. Fourth, as the pre-symptomatic share of infectious time increased, the effect of test frequency on the probability of an outbreak was strongly magnified largely independent of symptomatic self-isolation rates. Conclusions: Routine testing and isolation could significantly mitigate the risk of catastrophic PPP escapes, with the intervention's success varying based on pathogen characteristics. High shares of asymptomatic and pre-symptomatic transmission notably increased the relative risk reductions achieved by the intervention. These findings suggest prioritizing testing interventions for pathogens with high asymptomatic and pre-symptomatic transmission and highlight the symptomatic self-isolation rate as a policy intervention target.

17
Covariate adjustment for hierarchical outcomes and the win ratio: how to do it and is it worthwhile?

Hazewinkel, A.-D.; Gregson, J.; Bartlett, J. W.; Gasparyan, S. B.; Wright, D.; Pocock, S.

2026-03-31 cardiovascular medicine 10.64898/2026.03.30.26347966 medRxiv
Top 0.1%
18.7%
Show abstract

Objectives: Introducing a new covariate adjustment method for hierarchical outcomes using ordinal logistic regression, comparing it with existing approaches, and assessing whether adjustment improves power in randomized trials with hierarchical outcomes. Methods: We developed an ordinal regression-based method for covariate adjustment of the win ratio and compared it with three alternatives: probability index models, inverse probability weighting, and a randomization-based estimator. Methods were applied to the EMPEROR-Preserved rial and tested through extensive simulations involving two common hierarchical outcome structures: time-to-event composites, and composites combining time-to-event with quantitative measures. Simulations assessed impacts on estimates, standard errors, and power across prognostic and non-prognostic settings. Results: In RCT data and simulations, covariate adjustment consistently increased power when adjusting for prognostic baseline variables. Gains were comparable to or greater than those in conventional Cox models, with no power loss for non-prognostic covariates. Our ordinal approach performed similarly to existing methods while providing interpretable covariate effect estimates. Adjusting for baseline values of quantitative components yielded power gains according to the baseline-to-follow-up correlation. Conclusions: Covariate adjustment for prognostic variables meaningfully improves efficiency in win ratio analyses for hierarchical outcomes. Our ordinal method is easily implemented and facilitates covariate effect interpretation. We recommend the broader adoption of covariate adjustment and our ordinal method in randomized trials using hierarchical outcomes.

18
Testing out of quarantine

D'Agostino McGowan, L.; Lee, E. C.; Grantz, K. H.; Kucirka, L.; Gurley, E. S.; Lessler, J.

2021-02-01 infectious diseases 10.1101/2021.01.29.21250764 medRxiv
Top 0.1%
18.3%
Show abstract

Since SARS-CoV-2 emerged, a 14-day quarantine has been recommended based on COVID-19"s incubation period. Using an RT-PCR or rapid antigen test to "test out" of quarantine is a frequently proposed strategy to shorten duration without increasing risk. We calculated the probability that infected individuals test negative for SARS-CoV-2 on a particular day post-infection and remain symptom free for some period of time. We estimate that an infected individual has a 20.1% chance (95% CI 9.8-32.6) of testing RT-PCR negative on day five post-infection and remaining asymptomatic until day seven. We also show that the added information a test provides decreases as we move further from the test date, hence a less sensitive test that returns rapid results is often preferable to a more sensitive test with a delay.

19
Assessing the plausibility of subcritical transmission of 2019-nCoV in the United States

Blumberg, S.; Lietman, T. M.; Porco, T. C.

2020-02-11 epidemiology 10.1101/2020.02.08.20021311 medRxiv
Top 0.1%
18.1%
Show abstract

Rapid assessment of the transmission potential of an emerging or reemerging pathogen is a cornerstone of public health response. A simple approach is shown for using the number of disease introductions and secondary cases to determine whether the upper bound of the reproduction number exceeds the critical value of one.

20
Modeling the Effects of Routine Screening for Accidental Lab-Acquired Infections on the Risk of Potential Pandemic Pathogen Escape from High-Biosafety Research Facilities

Cohen, B.; Ren, B.; Hanage, W.; Menzies, N.; Croke, K.

2025-05-18 epidemiology 10.1101/2025.05.16.25327796 medRxiv
Top 0.1%
18.0%
Show abstract

Accidental lab-acquired infections (LAIs) risk releasing potential pandemic pathogens (PPPs) from BSL-3/4 facilities. We constructed a stochastic network infectious disease model to simulate how the probability of an outbreak of a pathogen resembling wild-type SARS-COV-2, following an initial LAI would be influenced by test-and-isolate interventions over a 100-day horizon. We varied test frequency (0-7 tests/week), peak sensitivity (50-100%), and isolation delay (0-3 days). For each of 192 parameter combinations, we conducted 1,000 simulations and used logistic regression to quantify how each parameter influenced the likelihood of an outbreak of 50 or more infections. Results indicated that even relatively infrequent routine testing significantly reduced the risk of outbreaks under diverse plausible scenarios, with greater reductions achieved at higher test frequencies. Once-weekly testing reduced outbreak risk by 52% under optimistic assumptions (80% sensitivity, 1-day delay) and by 29% under pessimistic assumptions (50% sensitivity, 2-day delay). Testing two and five times weekly yielded risk reductions of up to 62% and 71%, respectively, under optimistic assumptions, and 43% and 55%, respectively, under pessimistic assumptions. Logistic regression showed each additional weekly test decreased outbreak odds by 20%, each 10-point increase in test sensitivity reduced odds by 10%, and each additional isolation delay day increased odds by 15.5%. Interaction analyses revealed that longer isolation delays attenuated the protective effects of higher testing frequency and sensitivity. Routine lab-worker screening with prompt isolation substantially mitigates PPP escape risks. High-frequency testing has the greatest impact, and policymakers should consider implementing regular screening protocols.