Epidemiology
○ Ovid Technologies (Wolters Kluwer Health)
Preprints posted in the last 90 days, ranked by how well they match Epidemiology's content profile, based on 32 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.
Owusu-Boaitey, N.; Meyer, M. J.; Herrera-Esposito, D.; Bottcher, L.; Lukz, M.; Cook, S.; Stoto, M. A.; Kraemer, J. D.
Show abstract
Seroprevalence surveys reveal the extent of humoral immunity against pathogens such as severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), and under some circumstances represent cumulative incidence of prior infection. However, antibody waning - or seroreversion - biases these estimates by reducing assay sensitivity in a time-varying manner. Because assay sensitivity decays over time, naively using serosurveys can substantially bias estimates of SARS-CoV-2 cumulative incidence and fatality rates. The Bayesian assay-specific, time-varying sensitivity adjustment developed in this paper can reliably correct for this bias and account for the delay between infection and serosurvey. In seroprevalence studies conducted in the United States in 2020, adjusting for time-varying sensitivity increased cumulative incidence by up to 1.4-fold, with an adjustment of 1.08 for a national study. Our estimates contrast with a previously published 2-fold adjustment that did not account for assay design. This suggests that previous analyses overestimated cumulative incidence by applying seroreversion corrections that did not account for assay-specific effects, or underestimated cumulative incidence by not applying seroreversion corrections. These biases imply fatality rate underestimation and overestimation, respectively. Our model provides a framework for design-specific time-varying sensitivity corrections in seroprevalence surveys for other pathogens.
Yang, F.; Magee, A.; Morris, S. E.; Mathis, S. M.; Wiegand, R.; Iuliano, D. A.; Biggerstaff, M.; Olesen, S. W.
Show abstract
Vaccination can be a useful intervention for reducing infectious disease burden. Estimating numbers of vaccine-prevented health outcomes is one approach to quantifying the benefits of vaccination. Here we improve a method described by Foppa et al. (1) that assumes vaccination has only direct effects, that is, it cannot prevent infection or onward transmission of the disease. We rederive this method and derive an improved method that increases estimation accuracy with minimal additional analytical complexity. To evaluate the improved method, we simulated disease outbreaks and compared the accuracy of the two methods for estimating prevented disease outcomes. In 84% of simulations performed over a wide parameter space, the improved method had an equal or smaller estimation error compared to the original Foppa method, with 7.9-fold smaller mean error and 44-fold smaller standard deviation of errors. Our study improves a method for estimating prevented burden when assuming vaccination has only direct effects.
Qian, Y.; Song, Y.
Show abstract
Instrumental variable (IV) methods are widely used in health and social sciences to estimate causal treatment effects among compliers. In certain research settings, the instrument-treatment association (first stage) and the instrument-outcome association (reduced form) are each estimated from a different dataset. Two-Sample Instrumental Variables (TSIV), proposed by Angrist and Krueger (1992), addresses this by combining first-stage and reduced-form estimates from separate data sources into a single causal effect estimate. However, TSIV identification requires that instrument compliance behavior be consistent across the two samples, a condition that is rarely verified in practice. We show mathematically and empirically that when compliance differs between samples, the raw TSIV estimator does not converge to the true Local Average Treatment Effect (LATE) and instead attenuates toward a predictably biased limit proportional to the ratio of first-stage compliance rates between the two samples. To address this, we formalize a framework for estimating LATE with TSIV under two key assumptions: (1) Covariate Overlap, requiring that the two samples share sufficient common support in their covariate distributions, and (2) Compliance Transportability, requiring that compliance behavior is identical across populations after conditioning on observed covariates. We consider a setting in which a health policy instrument and outcomes are recorded in administrative claims while treatment and covariates are collected in a survey. We use a C-statistic derived from pooled covariates to detect population mismatch and an Inverse Probability Weighting (IPW) correction that reweights the first-stage sample to approximate the administrative covariate distribution. In Monte Carlo simulations across eight scenarios calibrated to a survey-Medicaid setting, IPW-TSIV reduces bias in estimating the LATE, achieving 88% reduction in the primary scenario, 82% under severe selection, and 79% when state-level expansion policy drives compliance heterogeneity. We further validate this framework using the Oregon Health Insurance Experiment, where partitioning the public-use lottery data (N = 24,646) into two non-overlapping samples with substantively meaningful compliance heterogeneity yields a verifiable benchmark against the true causal effect. IPW-TSIV reduces mean absolute bias by 71.6% relative to the oracle S2-specific LATE across 10 independent replications (C-statistic = 0.78), outperforms naive TSIV in all 10 splits, and reduces mean bias relative to the full-data LATE from +0.016 to +0.008. This framework provides applied researchers with actionable diagnostic thresholds to detect sample mismatch, validate transportability assumptions, and determine when structural TSIV estimation is reliable.
Wang, H.; Zhang, B.; Lei, Y.; Lu, Y.; Zhang, D.; Jian, X.; Zhu, Y.; Hu, W.; Chu, H.; Chen, Y.; Suchard, M. A.; Ryan, P. B.; Hripcsak, G.; Asch, D. A.; Lu, Y.; Bin, Y.; Schuemie, M. J.; Qiu, Y.; Chen, Y.
Show abstract
Glucagon-like peptide-1 receptor agonists (GLP-1RAs) have been linked to heterogeneous, potentially pleiotropic effects across organ systems, motivating outcome-wide comparative risk profiling in real-world data. A central challenge in such analyses is \emph{residual bias} that remains after adjustment for observed confounders, which can distort effect estimates and mis-calibrate uncertainty. We present distributional diagnosis and calibration (DC), which uses panels of negative control outcomes (NCOs) to diagnose residual bias and calibrate uncertainty. DC evaluates null behavior via $p$-value uniformity and empirical coverage across NCOs, and uses the empirical distribution of NCO effect estimates to calibrate confidence intervals for prespecified primary outcomes. DC is modular: it can wrap around commonly used causal inference methods and operates directly on summary statistics, supporting collaborative research under data-sharing constraints. Using electronic health records from a large U.S. clinical research network (152.7 million patients), we compared GLP-1RAs with sodium--glucose cotransporter~2 inhibitors across 15 prespecified outcomes spanning cardiovascular, mental health, and genitourinary domains using four causal estimators. Across outcomes and methods, DC diagnostics revealed substantial and method-dependent residual systematic error. DC calibration attenuated systematic error signals observed in negative controls and yielded more stable, better-calibrated estimates for clinical outcomes, supporting DC as a practical strategy to strengthen the credibility of real-world comparative effectiveness research.
Yechezkel, M.; Kapadia, B.; Hong, V.; Lim, J. T.; Reyes, I. A. C.; Pomichowski, M. E.; Davis, G. S.; Rodriguez Barraquer, I.; Mueller, N. F.; Tartof, S. Y.; Lewnard, J. A.
Show abstract
Importance: Doxycycline post-exposure prophylaxis (doxyPEP) is recommended to prevent bacterial sexually-transmitted infections (bSTIs) among certain men who have sex with men and transgender women. Impacts on bSTI transmission are unknown. Objective: To quantify the impact of doxyPEP implementation on bSTI risk, distinguishing indirect and direct protection. Design: Longitudinal cohort study, January 2022 to June 2025, with a nested test-negative design study among individuals potentially eligible to receive doxyPEP. Setting: Kaiser Permanente Southern California healthcare system. Participants: Individuals aged 16-59 years, including: 'at-risk males' assigned male sex at birth, who were living with HIV or who recently received HIV pre- or post-exposure prophylaxis; other males; and females. The nested test-negative design study included at-risk males who received bSTI testing after doxyPEP implementation. Exposures: Oral doxyPEP prescription dispense, defined as a 30+ dose supply of 200mg doxycycline absent accompanying bSTI diagnoses. Main outcomes and measures: Incident laboratory-confirmed gonorrhea, chlamydia, and syphilis. We quantified indirect effects as incidence rate ratios comparing observed to expected incidence of each bSTI, absent doxyPEP implementation, among doxyPEP non-recipients. We quantified direct effects among recipients via the adjusted odds ratio of recent doxyPEP dispenses among individuals testing positive or negative for each bSTI. Results: Analyses included 30,185 at-risk males, among whom 2,713 (9.0%) received 1+ doxyPEP fill; 1,212,854 other males; and 1,321,363 females. Among all at-risk males, overall doxycycline consumption increased 2.31-fold (95% confidence interval: 1.99-2.64; absolute increase by 18,953 [16,052-21,048] defined daily doses per 1,000 person-years) after doxyPEP implementation, without accompanying changes in consumption among other males or females. Resulting indirect protection was associated with 41.4% (95% confidence interval: 33.0-48.8%) and 29.0% (13.4-41.8%) lower-than-expected incidence of chlamydia and syphilis, respectively, among at-risk males who did not receive doxyPEP. We observed no indirect effect against gonorrhea among at-risk males, and no indirect effect against any bSTI among other males or females. Among 22,937 at-risk males in the nested test-negative design study, direct protection from doxyPEP was associated with 66.8% (47.6-78.7%) and 50.1% (2.1-74.2%) further reductions in chlamydia and syphilis risk, respectively, and no reduction in gonorrhea risk. Among doxyPEP recipients, one case of chlamydia and one case of syphilis was prevented for every 2,785 (1,553-5,491) and 20,001 (5,725-182,569) doses dispensed, respectively. Conclusions and relevance: DoxyPEP implementation conferred indirect as well as direct protection against chlamydia and syphilis among at-risk males. However, substantial volumes of antibiotic use were needed to realize this benefit. Tailoring doxyPEP guidance to circumstances associated with the greatest bSTI risk may be warranted to minimize unnecessary antibiotic use.
MA, Z.; XIANG, Y.; So, H.-C.
Show abstract
Abstract Purpose This study introduces a novel approach to address unmeasured confounding in terminal event studies using the prior event rate ratio (PERR) method. The proposed approach PERR_{proxy} used a proxy event to replace the original terminal event in the pre-exposure period, enabling the application of PERR in terminal event settings. Additionally, we also applied difference in difference (DID) regression, which is conceptually analogous to PERR to estimate the standard errors and confidence intervals of PERR_{proxy}. Methods We conducted numeric simulations to evaluate the validity of PERR_{proxy} approach and assessed its performance under varying levels of unmeasured confounding effects, baseline hazard ratios, and the correlation between the proxy and terminal events. To demonstrate its practical applicability, we also performed an empirical analysis to investigate the impact of severe hospitalized COVID-19 on circulatory system disease mortality using the PERR_{proxy}. Results In simulation studies, PERR_{proxy} effectively reduced the unmeasured confounding effects compared to the conventional methods. The performance of PERR_{proxy} was influenced by the strength of unmeasured confounding, baseline hazard ratios, and the correlation between the proxy and terminal outcomes. In addition, difference in difference (DID) regression had much faster computational speed for estimating standard errors and confidence intervals compared to bootstrap. In the empirical analysis, PERR_{proxy} identified that severe hospitalized COVID-19 as a significant risk factor for the circulatory system disease mortality and reduced the unmeasured confounding effects. Conclusions The PERR_{proxy} approach extends the applicability of the original PERR method to terminal event studies, offering a promising solution for addressing unmeasured confounding. Additionally, the DID regression framework provides a computationally efficient alternative for parameter estimation in PERR-based studies. However, careful consideration is still required in PERR_{proxy} for proxy events selection and other underlying assumptions of the PERR method to ensure valid results. Keywords: prior event rate ratio, unmeasured confounding, proxy event, terminal event study, observational study, electronic health records
Velasco Pardo, V.; Daines, L.; Katikireddi, S. V.; Ritchie, L.; Robertson, C.; Simpson, C. R.; McCowan, C.; Swallow, B.
Show abstract
Background During the COVID-19 pandemic, public health agencies used near real-time observational data to answer questions regarding vaccine effectiveness. However, traditional observational methods do not allow conclusions regarding counterfactual scenarios to be drawn from clinical data. Counterfactuals, which are outcomes that would have occurred under alternative interventions, can be used to formally assess the causal effects of public health interventions on health outcomes while accounting for the effects of confounding. Ideally individual patient data is used for the development of counterfactuals. Low-fidelity synthetic data may be useful for advancing methodological development where governance and privacy constraints prohibit access to sensitive personal data. Methods We simulated synthetic datasets based on the EAVE-II COVID-19 platform which has been limited to use for surveillance purposes. EAVE-II includes almost all resident people in Scotland registered with qualified general medical practitioners. Patient characteristics were simulated to reflect the known distribution of the Scottish population, accounting for dependencies between variables. Each synthetic dataset was encoded to different realistic scenarios for EAVEII 'ground truth' vaccine rollout and effectiveness results, explicitly stating the causal and confounding mechanisms, using a statistically sound method based on marginal structural models. Synthetic datasets of 100,000 individuals were then generated across five confounding scenarios and five severe outcome types. Results In scenarios with weak confounding, both unweighted and inverse probability of treatment weighted (IPTW) logistic regression recovered the true causal parameters. As confounding strength increased, only weighted models recovered the true mechanism. Conclusions Low-fidelity synthetic datasets simulated from EAVE-II data analysts to build and test causal inference pipelines, develop novel analysis pipelines, and train new researchers while awaiting access to real data. We showed how to generate synthetic datasets from a marginal structural model under different confounding scenarios.
Erly, B.; Raja, S.
Show abstract
Background. When a GLP-1 patient stops responding, should the clinician push the dose or change the drug? Observational answers conflate two distinct sources of confounding. Most early-week "escalations" in real-world data are FDA-mandated titration steps rather than deliberate clinical decisions, and patients who deviate do so for reasons we cannot observe. Semaglutide and tirzepatide are also not equivalent milligram-for-milligram, so naive class-switch comparisons mix mechanism and dose. We resolve both by restricting to post-titration patients and reclassifying treatments under the Whitley 2023 dose-equivalence framework. Methods. From 68,969 telehealth GLP-1 patients we built a post-titration cohort. Each patient's index time is the day they completed at least four weeks at therapeutic dose (Whitley tier 3 or higher: semaglutide 1.0 mg or tirzepatide 5 mg). Confirmed slow response is less than 5% total weight loss at the index, consistent with FDA weight-management drug-development guidance and AACE/ACE criteria. We compared four post-index strategies against continuing the current regimen: within-class dose escalation, equipotent class switch (a Whitley tier change of 1 or fewer), and class switch with potency increase. Direction-specific analyses split switches into semaglutide-to-tirzepatide and tirzepatide-to-semaglutide arms. Outcomes were percent weight loss at 12 and 24 weeks post-index. We estimated effects six ways: propensity-score matching; IPTW with linear and gradient-boosted propensities; the g-formula with linear and gradient-boosted outcome models; and AIPW, the doubly-robust estimator we use as the tiebreaker. Two-layer inverse probability of censoring weighting addressed strategy adherence and outcome ascertainment. We computed E-values, ran a negative-control specification, and stratified by tolerability. Results. The post-titration cohort comprised 24,876 confirmed slow responders. Within-class dose escalation produced a small consistent benefit at 24 weeks: AIPW +0.64 pp (95% CI +0.16 to +1.12), with five non-AIPW estimators ranging +0.47 to +0.76 pp. The continue arm itself lost an additional 8.07 pp over the same window (96% continued to lose), so escalation is a marginal addition to a substantial natural slope, not a rescue. Equipotent class switching from semaglutide to tirzepatide was inconclusive: linear and matching estimators ranged +1.26 to +1.80 pp, but AIPW was -0.33 pp (95% CI -1.27 to +0.60) with only 90 treated patients and limited propensity-score overlap. Class switch with simultaneous potency increase (sema to tirz) gave AIPW +0.65 pp at 12 weeks (95% CI +0.37 to +0.92, n = 80). A negative-control specification yielded ATE -0.13 pp, indicating the pipeline did not generate spurious signal. A held-out-fold prognostic-threshold sensitivity gave a null effect (+0.04 pp), correcting an earlier circular +1.06 pp estimate. Conclusions. Among confirmed post-titration slow responders, within-class dose escalation adds approximately 0.6 percentage points at 24 weeks on top of an 8 percentage point natural slope, consistently across six estimators including doubly-robust inference. This headline effect is small and not robust to modest unmeasured confounding (E-value 1.27) or to MNAR-style outcome attrition (tipping point delta approximately 1.2 pp), so it should be read as hypothesis-generating rather than practice-changing. Class-switching evidence is inconclusive; linear-estimator results suggesting benefit did not survive doubly-robust estimation in small treated samples with limited propensity overlap. The dose-ladder framework, with phase-specific evidence grading, is hypothesis-generating and insufficient on its own to change practice.
Habibdoust, A.; Song, X.
Show abstract
We applied the model-free GenericML framework using robust estimators (the Best Linear Predictor, Group Average Treatment Effects, and baseline profile classification) to reassess treatment effect heterogeneity (HTE) of intensive blood pressure (BP) using pooled data from two large-scale randomized control trials on BP controls. Among 10,712 participants with up to three years of follow-up, we estimated the differential risks of the primary cardiovascular outcome. Intensive BP control produced a modest average risk reduction (BLP b2 =-0.0136), but evidence for HTE was not statistically significant (BLP b2; p = 0.443; GATES top bottom contrast p = 0.471). Although CLAN grouped participants with metabolically adverse profiles into the highest predicted-benefit stratum, their treatment response did not differ significantly from that of the lowest-benefit group. Overall, while machine learning can identify clusters of high-risk baseline phenotypes, group-level inference revealed no meaningful HTE, underscoring the need for caution when interpreting individualized treatment predictions.
Mittelstaedt, R.; Helekal, D.; Kline, M. C.; Oliveira Roster, K. I.; Robbins, G. K.; Ard, K. L.; Grad, Y.
Show abstract
Background: Doxycycline post-exposure prophylaxis (doxy-PEP) reduces the incidence of bacterial sexually transmitted infections (STIs) among men who have sex with men and transgender women (MSMTW), but it may select for antimicrobial resistance (AMR). AMR development will depend in part how doxy-PEP changes rates of antibiotic use. Methods: We conducted a retrospective electronic medical record review of antibiotic prescriptions received by patients at the Massachusetts General Hospital Sexual Health Clinic from January 1, 2023, to December 27, 2025. Using a Bayesian negative binomial regression, we assessed the direct, indirect, and combined effects of doxy-PEP implementation on antibiotic prescription rates among doxy-PEP-eligible MSMTW who were receiving HIV pre-exposure prophylaxis. Results: Controlling for direct effects, the cohort's total antibiotic prescription rate decreased by 10% (0.90, 95% CI 0.87 - 0.94) for every 100 doxy-PEP starts. Doxy-PEP users were prescribed antibiotics at double the rate predicted in the absence of doxy-PEP implementation, and, controlling for indirect effects, received 3.40 (95% CI 2.95 - 3.91) times the antibiotic prescriptions of patients not using doxy-PEP. Doxy-PEP non-users were prescribed antibiotics at less than half the rate predicted in the absence of doxy-PEP. The full cohort's overall antibiotic prescription rate increased by 1.4 times after doxy-PEP implementation. Conclusions: Individuals who are taking doxy-PEP have higher antibiotic prescription rates, increasing selection for antibiotic-resistant bacteria in these individuals. However, doxy-PEP-driven decreases in the overall incidence of bacterial STIs have the potential to decrease selective pressure for resistant organisms in those not using doxy-PEP.
Creswell, R.; Golding, N.; Ryan, G. E.; Eales, O.; Price, D. J.; McCaw, J. M.; Shearer, F. M.
Show abstract
Knowledge of the true number of infections over time is valuable for accurately predicting the future course of an epidemic and planning effective interventions, but the number of cases reported offers only a noisy underestimate of the true number of infections. Disease surveillance strategies based on assessing subsets of the population for current infection (infection prevalence surveys) or antibody presence (seroprevalence surveys) yield crucial information about the number truly infected, but are expensive. To explore impact of survey design considerations--both sample size and sampling frequency--on inference of the number of incident infections over time, we coupled agent-based simulations of respiratory virus epidemics with simulations of infection prevalence and seroprevalence surveys. While returns diminish with increased sample size, we find inference generally improved by increasing survey frequency relative to participants-per-round for any given sample size. After survey rounds reach a sufficient frequency, comparable inference performance may be achieved with either more frequent rounds or more participants per round. Rolling designs with tests conducted each day tend to outperform designs in which testing is divided into discrete rounds. We also show that misspecified assumptions about seroreversion may substantially decrease the quality of inference results.
Otte, W. M.
Show abstract
Meta-analysis usually reduces each study to an effect estimate with a standard error and pools these by inverse-variance weighting: fixed effect (FE), random effects (RE), or unrestricted weighted least squares (UWLS). We propose information-geometric meta-integration (IGMI), representing each study by its sampling distribution, the Gaussian N(theta_i, Sigma_i), and pooling studies as a weighted Frechet mean (barycenter) under Bures-Wasserstein (BW), Fisher-Rao, or Wasserstein-Fisher-Rao (WFR) geometry. In the scalar fixed-variance case the BW barycenter mean is exactly the FE estimate; the minimized Frechet functional reproduces the Higgins-Thompson I^2 and DerSimonian-Laird tau^2 heterogeneity statistics; and a Frechet-scatter pivot reproduces the Hartung-Knapp-Sidik-Jonkman interval at m = 1 and yields an exact Hotelling F(m, K-m) region for m outcomes under proportional total covariances. WFR adds a robust outlier-resistant pool: as its length scale delta grows without bound it converges monotonically to BW, whereas finite delta gives a redescending M-estimator with rejection point exactly pi*delta. Simulations show calibrated multivariate coverage at small K, where Wald intervals undercover, and strong resistance of the equal-weight WFR pool to contamination. In 2,445 Cochrane meta-analyses, WFR most often wins leave-one-out predictive scoring. In 835 bivariate meta-analyses, the closed-form BW barycenter matches REML multivariate meta-analysis predictively and is exactly invariant to the unreported within-study correlation, unlike the likelihood estimate.
Mell, L. K.
Show abstract
In competing risks settings, covariate effects and group comparisons are usually assessed one event at a time - through log-rank or Cox tests on the cause-specific hazards, or Gray's test or Fine-Gray regression on a cumulative incidence function (CIF). This can obscure a clinically important quantity: the ratio between the event of interest and the competing event, since groups may differ little on the individual events yet differ sharply in their ratio. The generalized competing event (GCE) framework makes this ratio the object of inference; on the cause-specific scale the hazard ratio omega+(t) = lambda_1(t)/lambda_2(t) is estimated efficiently from a single stacked (Lunn-McNeil) model. We extend the framework to two scales that describe realized incidence. The subdistribution hazard ratio omega-tilde+(t) = lambda-tilde_1(t)/lambda-tilde_2(t) is estimated by a stacked, risk-set-weighted extension of the Lunn-McNeil construction; the cumulative-incidence ratio rho(t) = F_1(t)/F_2(t) - the odds that a subject's realized event by time t is the event of interest - by jackknife pseudo-observation regression of the Aalen-Johansen estimator. We relate the three contrasts: rho equals omega+ exactly under proportional cause-specific hazards, and equals omega-tilde+ only in the small-time limit under proportional subdistribution hazards, drifting toward 1 thereafter. The orthogonality that makes omega+ efficient is lost on both cumulative-incidence scales - omega tilde+ through overlapping weighted risk sets and shared censoring weights, rho through the shared all-cause survivor - so each carries a covariance term that must be handled and that bounds efficiency relative to the hazard-scale test. We derive the corresponding variances, study operating characteristics by simulation, illustrate on hypothetical prostate and head-and-neck cohorts, and provide an implementation in the gcemod R package.
Razzaghi, H.; Wieand, K.; Pinkney, A.; Bailey, C.
Show abstract
Research replication is essential to build trust in evidence produced from real-world data. However, methods for conducting and reporting these studies are lacking, particularly related to data quality and fitness assessments. We replicated a single-center study from Children's Hospital of Atlanta in a multi-institutional learning network (PEDSnet) to evaluate the long-term effects of hydroxyurea in children with severe sickle cell disease (SS/S{beta}0 genotype). An AS-IS arm applied the original study's criteria with no major data quality adjustments, while a Data Fitness Enhanced (DFE) arm used systematic data fitness assessment to inform adjustments to cohort inclusion criteria and variable definitions; both arms then replicated the original study's primary analyses. Data quality checks in the DFE arm refined cohort criteria and improved hydroxyurea capture, drug era computation, and hematology specialist mapping. The DFE cohort produced average treatment effects with higher face validity and greater concordance with the original study (e.g., change in ED visits: -0.44 (CI -0.60, -0.26) versus -0.36 (CI -0.57, -0.16) in the original study) than the AS-IS cohort (-0.08 (CI -0.26, 0.09)), which yielded several implausible results. These findings show that superficially plausible cohort characteristics do not guarantee valid results without transparent, systematic data fitness assessment.
Hamilton, F. W.; Ong, S. Y.; Swets, M.; Russell, C. D.; Underwood, J.
Show abstract
Background Staphylococcus aureus bacteraemia (SAB) is clinically heterogeneous. Potential heterogeneous treatment effects (HTE) have recently been identified through analysis of patient subgroups, identified using routine clinical variables.However, the impact of misclassifying patients into these groups is unclear, and practical strategies to improve HTE detection remain uncertain. Methods We performed a simulation study using published data from selected randomised trials and observational studies in SAB. We assessed the impact of varying classification accuracy (70%-100%) on i) power, ii) type I error, and iii) bias in post-hoc analyses of HTE. We then evaluated two strategies to improve performance: enrichment designs, in which only patients predicted to belong to a target subgroup are randomised, and the use of ordinal rather than binary outcomes. Results Even with perfect classification, post-hoc detection of heterogeneous treatment effects remained highly conditional on subgroup prevalence, baseline mortality, and effect size. One subgroup was detectable at moderate sample sizes; however, power was inadequate for all other subgroups even with sample sizes of 20,000. Decreasing classification accuracy reduced power, increased type I error, and introduced bias. Enrichment marginally improved power. Ordinal outcomes substantially improved performance when they matched the treatment-effect structure, but were worse when they did not. Conclusions Detecting HTE in SAB is challenging, but not uniformly infeasible. Feasibility depends on the interaction between subgroup frequency, baseline risk, classifier performance, and outcome choice. To advance stratified medicine in SAB, research should prioritize robust classifiers, outcome measures matched to the expected mechanism and pattern of treatment effect, and trial designs that acknowledge uncertainty in subgroup prevalence and treatment-effect structure.
Ma, Q.; Zhang, T.; Lin, D.
Show abstract
Abstract Objectives: To estimate surveillance-adjusted county-level residual syphilis risk, quantify posterior support for elevated risk, and identify the geographic distribution of stably high-risk areas across the contiguous United States and the District of Columbia. Methods: County-year primary and secondary syphilis counts from 3,109 counties during 2010-2022 were analyzed using a Bayesian negative-binomial spatial model with county-level covariates capturing social vulnerability and healthcare and surveillance related structure. Residual spatial risk, posterior exceedance probabilities, and stably high-risk counties were estimated. External validation examined whether county-level residual syphilis risk was associated with HIV and gonorrhea burden. Results: A total of 850 stably high-risk counties were identified. These counties were concentrated in the southeastern United States and along the Gulf Coast, with additional clusters in the north-central region and along the Atlantic and Pacific coasts. The social vulnerability index showed the strongest positive association with reported syphilis rates, followed by primary care physician density. External validation and sensitivity analyses showed that higher county-level residual syphilis risk estimates were positively associated with higher HIV diagnosis rates and gonorrhea rates, indicating that these estimates were not merely model-derived numerical outputs but were meaningfully related to the county-level distribution of sexually transmitted infection risk. These findings indicate that surveillance-adjusted residual spatial risk estimates and posterior exceedance probabilities may provide useful county-level evidence for syphilis control prioritization and resource allocation.
Chen, T.; Voorhies, K.; Reeson, A.; Seo, S.; Lee, S.; Hahn, G.; Hecker, J.; Prokopenko, D.; Hoth, K.; Kelly, R.; Lasky-Su, J. A.; Weiss, S.; Lange, C.; Lutz, S.
Show abstract
Mendelian Randomization (MR) is a popular tool for inferring causal relationships between traits using genetic variants as instrumental variables. These methods have been extended to also determine the direction of causality. However, causal direction cannot be inferred from a statistical test or estimation procedure (i.e. from data alone) without further assumptions and the methods operating characteristics and relative performances are not well understood. We conducted a comprehensive simulation study to illustrate this issue by evaluating type I error and power of 17 summary-based MR methods for inferring the effect direction. These methods fall within three methodological families: MR Steiger, Causal Direction (CD), and bidirectional MR approaches, with scenarios ranging across combinations of horizontal pleiotropy, unmeasured confounding, measurement error, longitudinal feedback, and varying sample sizes. While most methods achieved sufficient power levels under the alternative hypothesis in most scenarios, we found that every method was susceptible to inferring the wrong causal direction or under powered, and no method consistently maintained both correct type 1 error control and high power. In our applications, we evaluated the effect direction between the trait pairs body mass index (BMI) and major depressive disorder (MDD) and between BMI and asthma. To help researchers to evaluate the 17 methods to infer the effect direction and consider these challenges in their own data, we have developed MRdirection, an R package that runs the simulation studies examining the 17 directional MR methods across different user-defined scenarios. Our study, together with the accompanying R package, provides researchers with a tool for examining directional MR methods given different underlying assumptions.
Ng, S.-P.
Show abstract
The incidence rate ratio R is the standard measure for comparing event rates in clinical trials and epidemiology. In vaccine trials, the vaccine efficacy is VE = 1 - R. When events are rare, the two arm counts are Poisson. The estimator of R is heteroskedastic: its sampling variance changes with the data. So no fixed-width interval covers correctly everywhere. The usual log-Wald interval is undefined at zero events and covers poorly at small counts. Early vaccine and drug-safety readouts fall in exactly this regime. We show that a single reparameterization collapses this bivariate problem to an effective one-parameter family with a quadratic variance function, whose variance-stabilizing transformation is 2 arcsinh(sqrt(R)). The reduction yields a closed-form confidence interval for R. Its two leading errors, a curvature bias and the variability of the estimated scale, each admit a closed-form correction with no tuning constants. In a Monte Carlo study of our seven arcsinh variants and five competitors, the +Curve+Stu variant covers within 0.002 of the nominal 0.95 for about 50 control and 5 treatment events. Its width is on par with the best competitor. It avoids the conservatism and zero-count breakdown of log-Wald and MOVER. For moderate counts, we recommend this interval; for sparser data, our Bar-Lev and Enis count-shift variant is more robust. The result is a ready-to-use, closed-form interval for the low-count regime. We illustrate it on early Covid-19 vaccine-efficacy readouts and provide reference implementations in R and Python.
Marrec, L.; Lehtinen, S.
Show abstract
The rapid scale-up of HIV pre-exposure prophylaxis (PrEP) among men who have sex with men has coincided with rising rates of bacterial sexually transmitted infections (STIs), particularly Neisseria gonorrhoeae. This temporal association has raised concerns that PrEP may be driving a new STI epidemic. However, the epidemiological impact of PrEP reflects a trade-off between potential behavioral risk compensation, which increases transmission risk, and intensified clinical surveillance, which shortens infection duration. Determining whether PrEP amplifies or mitigates STI transmission therefore requires understanding how these competing effects balance at the population level. To address this question, we develop a transmission model stratified by sexual activity and PrEP use, derive simple analytical conditions governing changes in prevalence, incidence, and notification rates, and evaluate these dynamics using empirically informed parameter estimates. Our analysis demonstrates that current quarterly screening guidelines are generally sufficient to reduce both the true endemic prevalence and incidence of N. gonorrhoeae, successfully overcompensating for plausible reductions in condom use. We also confirm and expand on previous findings that clinical notification rates may surge even as the true disease burden declines, driven by the detection of previously undiagnosed asymptomatic infections. These findings suggest a shift in focus from the potential impact of PrEP on STI transmission to the consequences of increased STI diagnoses and treatment, particularly whether greater antibiotic consumption may accelerate the spread of antimicrobial resistance in N. gonorrhoeae.
Chervet, S.; Layan, M.; Boëlle, P.-Y.; Guedj, J.; van der Werf, S.; Kerneis, S.; Sermet-Gaudelus, I.; Cauchemez, S.; Opatowski, L.
Show abstract
Longitudinal household studies, combined with mathematical modeling, are widely used to characterize the drivers of respiratory pathogen transmission, including the effects of age and symptoms. In practice, household recruitment protocols vary across studies, potentially introducing biases into observed data. However, these biases are typically overlooked in statistical inference, and their impact on parameter estimates remains unknown. Here, we use synthetic household outbreak data simulated under different recruitment protocols to evaluate how recruiting through infected children affects estimates of age-specific infectiousness and susceptibility. We show that, under child-based recruitment, the standard likelihood, which accounts only for transmission dynamics, leads to underestimating child infectiousness and overestimating child susceptibility by more than 30%. We then propose a novel estimation framework that explicitly incorporates the household recruitment process into the likelihood and show that it substantially reduces these biases. Applying this new approach to a French household study conducted during the COVID-19 pandemic, we estimated that children under 6 had 49% lower infectiousness than teenagers and adults during the Alpha wave, whereas no difference was observed during the Omicron wave. This study demonstrates that ignoring recruitment protocols can bias key epidemiological parameter estimates and highlights the importance of accounting for study design.