Back

American Journal of Epidemiology

Oxford University Press (OUP)

All preprints, ranked by how well they match American Journal of Epidemiology's content profile, based on 67 papers previously published here. The average preprint has a 0.05% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Rescaling and Small Area Estimation of Health Survey Data as applied to Smoking Rates in Allegheny County, Pennsylvania

Stacy, S. L.; Chandra, H.; Gurewitsch, R.; Brink, L. L.; Robertson, L. B.; Wilson, D. O.; Yuan, J.-M.; Pyne, S.

2021-03-26 epidemiology 10.1101/2021.03.21.21254074 medRxiv
Top 0.1%
59.7%
Show abstract

We propose a novel, two-step method for rescaling health survey data and creating small area estimates of smoking rates using a Behavioral Risk Factor Surveillance System (BRFSS) survey administered in 2015 to participants living in Allegheny County, in the state of Pennsylvania, USA. The first step consisted of a spatial microsimulation to rescale location of survey respondents from zip codes to tracts based on census population distributions by age, sex, race, and education. The rescaling allowed us, in the second step, to utilize and select from available census tract specific ancillary data on social vulnerability for small area estimation (SAE) of local health risk using an area level version of a logistic linear mixed model. To demonstrate this new two-step algorithm, we estimated the ever-smoking rate for the census tracts of Allegheny County. The ever-smoking rate was slightly above 70% for two census tracts to the southeast of the city of Pittsburgh. Several tracts in the southern and eastern sections of Pittsburgh also had relatively high (>65%) ever-smoking rates. These small area estimates may be used in local public health efforts to target interventions and educational resources aimed at reducing cigarette smoking. Further, our new two-step methodology may be extended to small area estimation for other locations, and other health-related behaviors and outcomes.

2
Addressing spatial misalignment in population health research: a case study of US congressional district political metrics and county health data

Nethery, R. C.; Testa, C.; Tabb, L. P.; Hanage, W. P.; Chen, J. T.; Krieger, N.

2023-01-11 epidemiology 10.1101/2023.01.10.23284410 medRxiv
Top 0.1%
51.4%
Show abstract

Areal spatial misalignment, which occurs when data on multiple variables are collected using mismatched boundary definitions, is a ubiquitous obstacle to data analysis in public health and social science research. As one example, the emerging sub-field studying the links between political context and health in the United States faces significant spatial misalignment-related challenges, as the congressional districts (CDs) over which political metrics are measured and administrative units, e.g., counties, for which health data are typically released, have a complex misalignment structure. Standard population-weighted data realignment procedures can induce measurement error and invalidate inference, which has prompted the development of fully model-based approaches for analyzing spatially misaligned data. One such approach, atom-based regression models (ABRM), holds particular promise but has scarcely been used in practice due to the lack of appropriate software or examples of implementation. ABRM use "atoms", the areas created by intersecting all sets of units on which variables of interest are measured, as the units of analysis and build models for the atom-level data, treating the atom-level variables (generally unmeasured) as latent variables. In this paper, we demonstrate the feasibility and strengths of the ABRM in a case study of the association between political representatives voting behavior (CD-level) and COVID-19 mortality rates (county-level) in a post-vaccine period. The adjusted ABRM results suggest that more conservative voting record is associated with an increase in COVID-19 mortality rates, with estimated associations smaller in magnitude but consistent in direction with those of standard realignment methods. The results also indicate that ABRM may enable more robust confounding adjustment and more realistic uncertainty estimates, properly representing the uncertainties arising from all analytic procedures. We also implement the ABRM in modern optimized Bayesian computing programs and make our code publicly available, which may enable these methods to be more widely adopted.

3
Improving Assessment of Vaccine Effectiveness by Coupling Test-Negative Design Studies with Survival Models

Song, S.; Hitchings, M.; Yang, Y.; Longini, I.; N3C consortium,

2025-12-04 epidemiology 10.64898/2025.11.30.25341323 medRxiv
Top 0.1%
45.6%
Show abstract

The test-negative design (TND) has become a widely used observational study design for evaluating vaccine effectiveness, especially during the COVID-19 pandemic. Traditionally, TND has been viewed as a variant of the case-control study and largely limited to use with logistic regression models. In this paper, we first establish that TND can be framed as a special case of a cohort study, thereby opening the door to a wider range of analytical approaches. We then introduce the Prentice, Williams, and Peterson gap-time (PWP-GT) frailty model as a novel method for analyzing TND data, accounting for recurrent infections and time-dependent vaccination status. Through extensive simulation studies, we demonstrate that the proposed model outperforms conventional models commonly applied in TND-based vaccine effectiveness studies. Finally, we apply our method to data from the National COVID Cohort Collaborative, estimating the effectiveness of full and booster doses of Pfizers COVID-19 vaccines against both initial infection and reinfection during the Omicron variant circulation period in a real-world setting.

4
Multinational, Calibrated, Non-Laboratory Prevalent Disease Prediction and Survival Modeling for Diabetes, CKD, and CVD

Costa, A. M.; Badezet-Delory, I.

2025-12-11 epidemiology 10.64898/2025.12.10.25341981 medRxiv
Top 0.1%
45.2%
Show abstract

Reliable non-laboratory tools for assessment of probability of prevalent disease (PPD) are essential for scalable prevention, yet existing models are typically specific to single diseases, require laboratory tests, and show no or limited calibration across PPD strata, limiting scaling and public health utilization. We developed and validated a unified, non-invasive machine-learning model for simultaneous prediction of diabetes, chronic kidney disease, and cardiovascular disease PPD non-invasive predictors. The model was trained on 2011-2016 National Health and Nutrition Examination Survey data (n=29,903) and evaluated on an independent 2017-2020 test set (n=15,559). It demonstrated moderate-to-strong discrimination (C-statistic=0.80-0.90), stable precision-recall performance, and moderate to strong calibration (slope>0.94). Validation in an independent Korean population showed no or minimal degradation in discrimination and calibration performance, though more extensive validation is warranted. Predicted PPD was associated with cause-specific mortality over up to 7 years of follow-up, consistent with a predictor of latent disease burden. Each 10-percentage-point increase in predicted PPD was associated with roughly a two-fold higher hazard of disease-specific death (HR 2.00-2.20). We conclude that this model has potential as scalable, low-burden screening/surveillance aid, but note that it is not intended as a diagnostic or prognostic tool.

5
Regression with race-modifiers: towards equity and interpretability

Kowal, D. R.

2024-01-04 epidemiology 10.1101/2024.01.04.23300033 medRxiv
Top 0.1%
45.0%
Show abstract

The pervasive effects of structural racism and racial discrimination are well-established and offer strong evidence that the effects of many important variables on health and life outcomes vary by race. Alarmingly, standard practices for statistical regression analysis introduce racial biases into the estimation and presentation of these race-modified effects. We advocate abundance-based constraints (ABCs) to eliminate these racial biases. ABCs offer a remarkable invariance property: estimates and inference for main effects are nearly unchanged by the inclusion of race-modifiers. Thus, quantitative researchers can estimate race-specific effects "for free"--without sacrificing parameter interpretability, equitability, or statistical efficiency. The benefits extend to prominent statistical learning techniques, especially regularization and selection. We leverage these tools to estimate the joint effects of environmental, social, and other factors on 4th end-of-grade readings scores for students in North Carolina (n = 27, 638) and identify race-modified effects for racial (residential) isolation, PM2.5 exposure, and mothers age at birth.

6
Variations in the results of nutritional epidemiology studies due to analytic flexibility: Application of specification curve analysis to red meat and all-cause mortality

Wang, Y.; Pitre, T.; Wallach, J. D.; de Souza, R. J.; Jassal, T.; Bier, D.; Patel, C. J.; Zeraatkar, D.

2023-12-21 epidemiology 10.1101/2023.12.19.23300248 medRxiv
Top 0.1%
42.1%
Show abstract

ObjectiveTo present an application of specification curve analysis--a novel analytic method that involves defining and implementing all plausible and valid analytic approaches for addressing a research question--to nutritional epidemiology. Data sourceNational Health and Nutrition Examination Survey (NHANES) 2007 to 2014 linked with National Death Index. MethodsWe reviewed all observational studies addressing the effect of red meat on all-cause mortality, sourced from a published systematic review, and documented variations in analytic methods (e.g., choice of model, covariates, etc.). We enumerated all defensible combinations of analytic choices to produce a comprehensive list of all the ways in which the data may reasonably be analyzed. We applied specification curve analysis to NHANES data to investigate the effect of unprocessed red meat on all-cause mortality, using all reasonable analytic specifications. ResultsAmong 15 publications reporting on 24 cohorts included in the systematic review on red meat and all-cause mortality, we identified 70 unique analytic methods, each including different analytic models, covariates, and operationalizations of red meat (e.g., continuous vs. quantiles). We applied specification curve analysis to NHANES, including 10,661 participants. Our specification curve analysis included 1,208 unique analytic specifications. Of 1,208 specifications, 435 (36.0%) yielded a hazard ratio equal to or above 1 for the effect of red meat on all-cause mortality and 773 (64.0%) below 1, with a median hazard ratio of 0.94 [IQR: 0.83 to 1.05]. Forty-eight specifications (3.97%) were statistically significant, 40 of which indicated unprocessed red meat to reduce all-cause mortality and 8 of which indicated red meat to increase mortality. ConclusionWe show that the application of specification curve analysis to nutritional epidemiology is feasible and presents an innovative solution to analytic flexibility. LimitationsAlternative analytic specifications may address slightly different questions and investigators may disagree about justifiable analytic approaches. Further, specification curve analysis is time and resource-intensive and may not always be feasible.

7
An enhanced method for calculating trends in infections caused by pathogens transmitted commonly through food

Weller, D.; Ray, L.; Payne, D. C.; Griffin, P. M.; Hoekstra, R. M.; Rose, E. B.; Bruce, B. B.

2022-09-17 epidemiology 10.1101/2022.09.14.22279742 medRxiv
Top 0.1%
40.9%
Show abstract

This brief methods paper is being published concomitantly with "Preliminary Incidence and Trends of Infections Caused by Pathogens Transmitted Commonly Through Food-- Foodborne Diseases Active Surveillance Network, 10 U.S. Sites, 2016-2021" in Morbidity and Mortality Weekly Reports (MMWR). That article describes the application of the new model described here to analyze trends and evaluate progress towards the prevention of infection from enteric pathogens in the United States.

8
Constructing and analyzing a synthetic life course cohort based on pooling two data sources: A case study of early adulthood depression symptomatology and late-life cognition

Zimmerman, S. C.; Buto, P.; Kezios, K.; Zeki Al Hazzouri, A.; Glymour, M. M.

2026-02-27 epidemiology 10.64898/2026.02.25.26347113 medRxiv
Top 0.1%
39.7%
Show abstract

BackgroundSynthetic cohorts created by combining two cohorts can be useful when no single data set includes both the exposure and outcome data of interest. We estimate the effects of depression in early adulthood on later-life memory outcome using two nationally representative cohorts separately and in a synthetic sample. MethodsWe used the National Longitudinal Study of Youth 1979 (NLSY; N=5,747) and the Health and Retirement Study (HRS; N=6,846) and a synthetic cohort combining exposure data from N=5,680 NLSY participants (born 1957-1965) aged 55-63 in 2020 who completed midlife cognitive assessment between 2006-2020 with outcome data from N=9,726 HRS participants born 1957-1964 who completed cognitive assessments when 47-63 years old and every 2-years thereafter. A 6-item version of the Centers for Epidemiologic Studies-Depression (CES-D) score (range 0-6) was measured from late adolescence through midlife in NLSY and in midlife in HRS. Memory was measured as the sum of immediate and delayed word recall scores up to twice in NLSY at age 48+ and up to 10 times in HRS at age 50+. We generated a synthetic life course cohort, matching HRS participants to NLSY participants based on 10 variables measured in midlife in both cohorts and posited to either confound or mediate the association between early life depressive symptoms and late-life memory. Matching variables included midlife depression and memory. We used confounder-adjusted linear mixed models to estimate the association between earliest reported depressive symptoms in NLSY and HRS with memory in the respective data sets and evaluated associations of early life depression symptoms with the repeated later life memory measures in the synthetic cohort. ResultsIn NLSY, each increment in CES-D at age 23-31 was associated with lower average memory scores ({beta}NLSY_level=-0.050 95%CI (-0.097,-0.003)) in midlife but no detectable difference in rate of memory decline ({beta}NLSY_slope=-0.070 95%CI (-0.382,0.242). In HRS, CES-D at average age 53 was associated with lower average memory ({beta}HRS_level=-0.163 (-0.199, -0.128)) but not rate of decline ({beta}HRS_slope=-0.021 (-0.062, 0.020)). In the synthetic cohort, CES-D at age 23-27 was associated with lower memory score at age 50+ ({beta}synth_level=-0.044 95%CI (-0.085,-0.003)) but not associated with rate of cognitive decline ({beta}synth_slope=0.005 95%CI (-0.052,0.062)). ConclusionsDepressive symptoms ages 23-31 predicted mid- to late-life memory function but had no clear association with memory decline. Combining data across cohorts spanning separate, but overlapping, parts of the life course is a promising approach to overcome data limitations in life course research, but it requires careful implementation to ensure that assumptions are met and estimates are appropriately interpreted.

9
Identifying US Counties with High Cumulative COVID-19 Burden and Their Characteristics

Li, D.; Gaynor, S. M.; Quick, C.; Chen, J. T.; Stephenson, B. J. K.; Coull, B. A.; Lin, X.

2021-01-12 epidemiology 10.1101/2020.12.02.20234989 medRxiv
Top 0.1%
38.9%
Show abstract

Identifying areas with high COVID-19 burden and their characteristics can help improve vaccine distribution and uptake, reduce burdens on health care systems, and allow for better allocation of public health intervention resources. Synthesizing data from various government and nonprofit institutions of 3,142 United States (US) counties as of 12/21/2020, we studied county-level characteristics that are associated with cumulative case and death rates using regression analyses. Our results showed counties that are more rural, counties with more White/non-White segregation, and counties with higher percentages of people of color, in poverty, with no high school diploma, and with medical comorbidities such as diabetes and hypertension are associated with higher cumulative COVID-19 case and death rates. We identify the hardest hit counties in US using model-estimated case and death rates, which provide more reliable estimates of cumulative COVID-19 burdens than those using raw observed county-specific rates. Identification of counties with high disease burdens and understanding the characteristics of these counties can help inform policies to improve vaccine distribution, deployment and uptake, prevent overwhelming health care systems, and enhance testing access, personal protection equipment access, and other resource allocation efforts, all of which can help save more lives for vulnerable communities. Significance statementWe found counties that are more rural, counties with more White/non-White segregation, and counties with higher percentages of people of color, in poverty, with no high school diploma, and with medical comorbidities such as diabetes and hypertension are associated with higher cumulative COVID-19 case and death rates. We also identified individual counties with high cumulative COVID-19 burden. Identification of counties with high disease burdens and understanding the characteristics of these counties can help inform policies to improve vaccine distribution, deployment and uptake, prevent overwhelming health care systems, and enhance testing access, personal protection equipment access, and other resource allocation efforts, all of which can help save more lives for vulnerable communities.

10
Mechanism Matters: A Monte Carlo Evaluation of Estimator Validity and Collider Bias in Environmental Mixture Epidemiology

Obeng-Gyasi, E.

2026-05-26 epidemiology 10.64898/2026.05.25.26354044 medRxiv
Top 0.1%
36.2%
Show abstract

Background: Mixture epidemiology deploys sophisticated estimators, Bayesian kernel machine regression with causal mediation analysis (BKMR-CMA), quantile G-computation (QGC), and parametric G-computation, alongside conventional regression. Comparative evaluations have assumed additive, non-mediated data-generating processes, leaving conditions under which estimator choice determines causal validity uncharacterized. Methods: We developed a simulation framework using military-relevant exposure distributions (metals, per- and polyfluoroalkyl substances [PFAS], polychlorinated biphenyls [PCBs]) and allostatic load (AL) across three deployment tiers, with parameters drawn from military occupational health and contamination literature. Four data-generating processes were specified as directed acyclic graphs: direct effects with confounding (M1), full mediation through AL (M2), synergistic AL-exposure interaction (M3), and collider structure (M4). We evaluated ordinary least squares (OLS), QGC, G-computation, and BKMR-CMA on bias, root mean squared error, and 95% confidence interval coverage across 500 Monte Carlo replications at n = 500 and n = 1,000. Results: No estimator dominated across all mechanisms. Under M1, OLS and G-computation produced near-identical modest positive bias; BKMR-CMA achieved lower root mean squared error through kernel shrinkage. Under M2, BKMR-CMA exhibited severe positive bias for AL (mean bias = +0.579 SD units; coverage = 32.8%). Under M3, BKMR-CMA was the only estimator achieving nominal 95% coverage for AL (95.2%), while regression-based approaches fell to 83.6%. Under M4, G-computation produced persistent bias and near-zero coverage for lead, reflecting structural non-identification. Conclusions: Estimator validity is fundamentally mechanism-dependent. Researchers should base estimator choice on explicit causal assumptions about whether AL functions as confounder, mediator, moderator, or collider, particularly in military and occupational cohorts. We provide a mechanism-to-estimator mapping for applied researchers.

11
A roadmap to account for reporting delays for public health situational awareness: a case study with COVID-19 and dengue in United States jurisdictions.

Lopez, V. K.; Bastos, L. S.; Codeco, C.; Johansson, M. A.

2024-11-13 epidemiology 10.1101/2024.11.09.24315999 medRxiv
Top 0.1%
35.4%
Show abstract

BackgroundDecision-making in public health is limited by data availability where the most recent reports do not reflect the actual trajectory of an epidemic. Nowcasting is a modeling tool that can estimate eventual case counts by accounting for reporting delays. While these tools have generated reliable predictions when designed for specific use cases, several limitations exist when scaling the models to systems composed of multiple distinct surveillance systems. We seek to identify flexible application of nowcasting models to address these problems. MethodsWe used a previously developed Bayesian nowcasting tool, which dynamically estimates delay probabilities up to a user-defined maximum delay using a user-defined training window. We tested automated approaches to select the maximum delay and training window, setting maximum delay values at the 90th, 95th, and 99th quantile distribution of the most recently reported data and training windows to the maximum delay plus one week or multiplied by 1.5 or 2.0. We generated and evaluated 321 nowcasts for COVID-19 cases in six U.S. states and dengue cases in Puerto Rico. We assessed prediction error and precision via logarithmic scoring and coverage metrics for the most recent three weeks of predictions in each nowcast. We used these metrics to further assess why nowcasts may fail and to compare predictions generated from three different publicly available tools. ResultsUsing recent data to estimate dynamic delay and training window parameters resulted in nowcast with less error relative to nowcasts made with static parameters for long historic periods. Nowcasts likely to fail could be predicted a priori by the relative width of the prediction intervals and the permutation entropy of the epidemic trend. More complex models do not necessarily improve performance, for example, a model with random effects for reporting periodicity did not improve nowcasts compared to a simple model which fit the observed epidemic trend. ConclusionsWe tested multiple systems for scaling up nowcasts in a flexible framework. We recommend using dynamic parameter selection and creating a system to suppress nowcasts likely to fail. This requires collaboration with surveillance colleagues to implement data-driven choices to improve the utility of predictions for decision-making.

12
its2s: a Python package for two-stage interrupted time series analysis using machine learning

Wilner, L.; Casey, J. A.; Mooney, S. J.; Do, V.; Ma, Y.; Benmarhnia, T.; Dey, A. K.

2026-07-06 epidemiology 10.64898/2026.07.02.26357175 medRxiv
Top 0.1%
34.6%
Show abstract

When randomized controlled trials are infeasible, researchers may leverage natural experiments for causal inference. Interrupted time-series (ITS) designs compare observed post-event trends to counterfactual predictions from pre-event data. Two-stage ITS designs use flexible models to generate optimized counterfactual predictions in the first stage, then estimate intervention effects by comparing observed to predicted outcomes in the second stage. Fitting high-dimensional versions of these models is challenging, requiring systematic infrastructure to ensure rigor and reproducibility. In response, we developed its2s, an open-source Python package implementing the two-stage ITS design with machine learning. its2s allows users to specify an intervention date and training/testing periods, select among built-in model architectures (e.g., Prophet-XGBoost, NeuralProphet), and generate confidence intervals via moving block bootstrap, preserving temporal autocorrelation in residuals. its2s layers defaults, configuration files, and runtime overrides to support workflows ranging from rapid default implementations to highly tailored analyses. We validated its2s using two case studies: a simulation with a 12% policy effect, recovering the true effect as 11.77%, and an analysis of the 2021 Pacific Northwest heat dome, finding 53% excess injury mortality over the following three weeks. its2s provides a flexible, reproducible framework for ITS-based quasi-experimental research, lowering barriers to rigorous machine learning-based counterfactual modeling.

13
Application of the Adaptive Validation Design to estimate the association between transmasculine/transfeminine status and self-inflicted injury among transgender and gender-nonconforming children and adolescents.

Collin, L. J.; MacLehose, R. F.; Ahern, T. P.; Goodman, M.; Lash, T. L.

2020-02-23 epidemiology 10.1101/2020.02.20.20024182 medRxiv
Top 0.1%
34.1%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWO_ST_ABSBackgroundC_ST_ABSAn internal validation substudy compares an imperfect measurement of a variable with a gold standard measurement in a subset of the study population. Validation data permit calculation of a bias-adjusted estimate, expected to equal the association that would have been observed had the gold standard measurement been available for the entire study population. Guidance on optimal sampling of participants to include in validation substudies has not considered monitoring validation data as they accrue. In this paper, we develop and apply the framework of Bayesian monitoring to determine when sufficient validation data have been collected to yield a bias-adjusted estimate of association with a prespecified level of precision. MethodsWe demonstrate the utility of this method using the Study of Transition, Outcomes and Gender--a cohort study of transgender and gender non-conforming children and adolescents. Transmasculine and transfeminine status were determined from the gender code in the electronic medical record at cohort enrollment. This status is known to be misclassified because it can indicate either gender identity or sex recorded at birth. Our interest is in the association between transmasculine and transfeminine status and self-inflicted injury. To address possible exposure misclassification, we demonstrate the methods ability to determine when sufficient validation data have been collected to calculate a bias-adjusted estimate of association that is less than 80% greater than the precision of the conventional estimate. ResultsIn the conventional age-adjusted analysis, we observed that transmasculine children and adolescents were 1.80-fold more likely to inflict self-harm than transfeminine youths (95%CI 1.27, 2.55). Using the adaptive validation approach, 200 cohort members were required for validation to yield a bias-adjusted estimate of OR=3.03 (95%CI 1.76, 5.56), which was similar to the bias-adjusted estimate using complete validation data (OR=2.63, 95%CI 1.67, 4.23). ConclusionsOur method provides a novel approach to effective and efficient estimation of classification parameters as validation data accrue. This method can be applied within the context of any parent epidemiologic study design, and modified to meet alternative criteria given specific study or validation study objectives.

14
Bias from small-count suppression in county-level cancer disparity estimates: a calibrated simulation study

gahan, k.

2026-06-08 epidemiology 10.64898/2026.06.05.26355021 medRxiv
Top 0.1%
33.8%
Show abstract

Abstract Background. Area-level cancer disparities are routinely estimated from public county data in which rates based on small counts (fewer than 16 cases or deaths) are suppressed. Analysts typically drop suppressed counties (complete-case analysis). Because suppression depends on case counts tied to population size and demographic composition, this missingness may be informative, but its effect on the disparity estimate has not, to our knowledge, been quantified. Methods. In a cross-sectional ecological study of 3,143 U.S. counties (analytic sample 3,018 with computable exposure) using one frozen public release of NCI State Cancer Profiles incidence and mortality data and ACS 2018-2022 5-year data, we estimated the most- versus least-deprived ICE(race+income) quintile rate ratio (RR) and rate difference for female breast, stomach, and cervix cancers under four suppression-handling methods: complete-case, available-case, bounding, and model-based small-area estimation. We characterized which counties were erased, and, following the ADEMP framework, ran a Monte Carlo simulation (1,000 replicates per cell; Monte Carlo standard error of bias approximately 0.0025) calibrated to the release to measure bias against a known truth. Analyses were pre-registered. Results. The suppressed fraction rose with rarity: 7.4% of counties for breast, 61.3% for stomach, and 75.7% for cervix incidence. Suppression was concentrated in the most-deprived quintile (cervix, 81.8% suppressed vs 63.8% least-deprived) and overwhelmingly removed rural rather than minority residents (cervix: 81% of the rural but 9% of the minority population erased). For breast (little suppression) the RR was 0.87 (95% CI 0.85-0.89) and identical across methods; for cervix incidence the complete-case RR (1.56) exceeded the model-based estimate (1.50), and for cervix mortality (91% suppressed) complete-case (1.86) exceeded model-based (1.56) by 16% with a wide bounding interval (1.88-2.62). In calibrated simulation, population-weighted complete-case bias was small (less than 2%) at the observed deprivation-county-size correlation and grew with rarity, threshold, and unweighted aggregation; its direction was conditional, becoming positive (over-estimation) as deprived counties became smaller. Conclusions. Complete-case handling of suppressed counties over-estimates rare-cancer area disparities relative to methods that retain them, while silently erasing most of the rural and most-deprived communities the estimate is meant to represent. The effect is negligible for common cancers and grows with rarity. Public-data disparity analyses should report the suppressed fraction and use bounded or model-based estimates by default. Keywords: cancer disparities; small-count suppression; Index of Concentration at the Extremes; informative missingness; small-area estimation; rural health.

15
Sequence Analysis as an approach to characterize variables that unfold over time: implementation and practical considerations for epidemiologists

Pacca, L.; Dang, K. V.; Koenig, L. R.; Duarte, C. d.; Gaye, S. A.; Harrati, A.; Vable, A. M.

2024-06-19 epidemiology 10.1101/2024.06.18.24308957 medRxiv
Top 0.1%
32.2%
Show abstract

Characterizing longitudinal trajectories of social exposures or health outcomes is a persistent challenge, but can be accomplished with sequence analysis, a data-driven approach that can differentiate timing, order and duration of events. We present practical guidance on implementing sequence analysis for epidemiologists with the goal of providing clear advice on decision points and tradeoffs. We Introduce the three main steps of sequence analysis: (1) coding longitudinal processes as trajectories of ordered events for a set of individuals, (2) measuring dissimilarity between individual trajectories, and (3) performing cluster analysis to group similar trajectories. Each of these steps presents researchers with several decision points, such as data cleaning rules, options for evaluating sequence dissimilarity, and choices of clustering algorithms to group trajectories. After outlining each of the sequence analysis steps, we provide an applied example of sequence analysis in which we create and group transition-to-retirement trajectories from age 51-75 for a sample of 9,189 Health and Retirement Study participants using self-reported employment information, then estimate the association between transition-to-retirement groups and self-rated health. Our paper seeks to guide epidemiologists through the analytic decisions and implementation challenges of sequence analysis as this approach is increasingly implemented and undergoes methodological advances.

16
Modeling suicide mortality in US counties using population socioeconomic indicators

Kandula, S.; Martinez-Ales, G.; Rutherford, C.; Gimbrone, C.; Olfson, M.; Gould, M. S.; Keyes, K. M.; Shaman, J.

2022-06-06 public and global health 10.1101/2022.06.06.22275887 medRxiv
Top 0.1%
31.3%
Show abstract

BackgroundSuicide is one of the leading causes of death in the United States and population risk prediction models can inform the type, location, and timing of public health interventions. Here, we report the development of a prediction model of suicide risk using population characteristics. MethodsAll suicide deaths reported to the Nation Vital Statistics System between 2005-2019 were identified, and age, sex, race, and county-of-residence of the decedents were extracted to calculate baseline risk. County-wise annual measures of socioeconomic predictors of suicide risk -- unemployment, weekly wage, poverty prevalence, median household income, and population density -- along with two state-wise measures of prevalence of major depressive disorder and firearm ownership were compiled from public sources. Conditional autoregressive (CAR) models, which account for spatiotemporal autocorrelation in response and predictors, were used to estimate county-level risk. ResultsEstimates derived from CAR models were more accurate than from models not adjusted for spatiotemporal autocorrelation. Inclusion of suicide risk/protective covariates further reduced errors. Suicide risk was estimated to increase with each standard deviation increase in firearm ownership (2.8%), prevalence of major depressive episode (1%) and unemployment (2.8%). Conversely, risk was estimated to decrease by 4.3% for each standard deviation increase in both median household income and population density. Increased heterogeneity of risk across counties was also noted. ConclusionsArea-level characteristics and the CAR model structure can estimate population-level suicide risk and thus inform decisions on resource allocation and focused interventions during outbreaks.

17
A Simulation Study Comparing Multiple Imputation and Complete Case Analysis for Handling Missing Preschool Body Mass Index

Savu, A.; Dover, D. C.; Hajihosseini, M.; Gaudet, L. A.; Kaul, P.

2026-08-14 epidemiology 10.64898/2026.08.13.26360115 medRxiv
Top 0.1%
31.3%
Show abstract

Background and Objective. Missing data frequently occurs in health databases and can bias analyses if not correctly dealt with. Using real-world data, we compared complete-case and multiple-imputation methods for recovering true parameters of a multivariable logistic regression model for the association between maternal glucose levels during pregnancy and child excess weight at preschool age, where missing values were present in as much as 30% of our sample. Methods. This study utilized a cohort of 130,424 children with complete preschool-age body mass index (BMI) measurements from the Calgary and Edmonton health regions of Alberta, Canada. In the complete BMI data, we introduced missingness through deletion following three distinct mechanisms: missing completely at random (MCAR), at random (MAR), and not at random (MNAR). To handle the missing data created, we employed complete-case and multiple-imputation methods. Maternal glucose levels during pregnancy were categorized into five groups and its association with child excess weight at pre-school age was determined based on a logistic regression model using the full observed data (yielding true values), observed data that was not deleted (complete-case estimates), and imputed data (multiple-imputation estimates). The accuracy of complete-case and multiple-imputation estimates were evaluated against the true values. Finally, we conducted a sensitivity analysis for the MNAR mechanism using pattern-mixture models with an additive shift. Results. Under MCAR and MAR, multiple-imputation generally outperformed complete-case, yielding smaller absolute and relative bias. Both methods achieved high significance ([≥] 0.96) for most effects. Mean squared errors for multiple-imputation and complete-case were similar missing completely at random, missing at random, and coverage was consistently high ([≥] 0.99). Under MNAR, both complete-case and multiple-imputation showed poor performance regarding bias and statistical significance. Sensitivity analysis using pattern-mixture models indicated performance varied by specific effect. Conclusions. Under MCAR and MAR, multiple-imputation introduced higher bias but demonstrated superior overall performance based on mean squared error and restored statistical power. Conversely, both methods failed under MNAR, where pattern-mixture modeling sensitivity analyses revealed highly variable, effect-specific performance due to unverifiable shift assumptions. When faced with missing data, researchers should assess missingness mechanisms, report both complete-case and multiple-imputation estimates under MCAR/MAR while accounting for power-versus-bias tradeoffs, and employ pattern-mixture sensitivity analyses to test robustness when MNAR is plausible.

18
A two-step penalization and shrinkage approach for binary response data that is jointly separated and correlated: The effects of social networks on diarrheal disease

Hegde, S.; Eisenberg, J. N.; Beesley, L. J.; Mukherjee, B.

2024-03-18 epidemiology 10.1101/2024.03.13.24304191 medRxiv
Top 0.1%
29.0%
Show abstract

Epidemiologic data often violate common modeling assumptions of independence between subjects due to study design. Statistical separation is also common, particularly in the study of rare binary outcomes. Statistical separation for binary outcomes occurs when regions of the covariate space have no variation in the outcome, and separation can negatively impact the validity of logistic regression model parameters. When data are correlated, we generally use multi-level modeling for parameter estimation, and statistical approached have also been developed for handling statistical separation. Approaches for analyzing data with both separation and complex correlation, however, are not well-known. Extending prior work, we demonstrate a two-stage Bayesian modeling approach to account for both separated and highly correlated data through a motivating example examining the effect of social ties on Acute Gastrointestinal Illness (AGI) in rural Ecuador. The two-stage approach involves fitting a Bayesian hierarchical model to account for correlation using priors derived from parameter estimates from a Firth-corrected logistic regression model to account for separation. We compare estimates from the two-stage approach to standard regression methods that only account for either separation or correlation. Our results demonstrate that correctly accounting for separation and correlation when both are present can potentially provide better inference.

19
Predicting Long COVID in the National COVID Cohort Collaborative Using Super Learner

Butzin-Dozier, Z.; Ji, Y.; Li, H.; Coyle, J.; Shi, J.; Phillips, R. V.; Mertens, A.; Pirracchio, R.; van der Laan, M. J.; Patel, R. C.; Colford, J. M.; Hubbard, A. E.; on behalf of the National COVID Cohort Collaborative (N3C) Consortium*,

2023-08-04 epidemiology 10.1101/2023.07.27.23293272 medRxiv
Top 0.1%
28.1%
Show abstract

Post-acute Sequelae of COVID-19 (PASC), also known as Long COVID, is a broad grouping of a range of long-term symptoms following acute COVID-19 infection. An understanding of characteristics that are predictive of future PASC is valuable, as this can inform the identification of high-risk individuals and future preventative efforts. However, current knowledge regarding PASC risk factors is limited. Using a sample of 55,257 participants from the National COVID Cohort Collaborative, as part of the NIH Long COVID Computational Challenge, we sought to predict individual risk of PASC diagnosis from a curated set of clinically informed covariates. We predicted individual PASC status, given covariate information, using Super Learner (an ensemble machine learning algorithm also known as stacking) to learn the optimal, AUC-maximizing combination of gradient boosting and random forest algorithms. We were able to predict individual PASC diagnoses accurately (AUC 0.947). Temporally, we found that baseline characteristics were most predictive of future PASC diagnosis, compared with characteristics immediately before, during, or after COVID-19 infection. This finding supports the hypothesis that clinicians may be able to accurately assess the risk of PASC in patients prior to acute COVID diagnosis, which could improve early interventions and preventive care. We found that medical utilization, demographics, anthropometry, and respiratory factors were most predictive of PASC diagnosis. This highlights the importance of respiratory characteristics in PASC risk assessment. The methods outlined here provide an open-source, applied example of using Super Learner to predict PASC status using electronic health record data, which can be replicated across a variety of settings.

20
Biomarker Panel Selection Explains Heterogeneity in Allostatic Load-Mortality Risk: A Specification-Curve Analysis

Patel, P. C.

2026-05-08 epidemiology 10.64898/2026.05.06.26352579 medRxiv
Top 0.1%
27.2%
Show abstract

The allostatic load (AL)-mortality association is well-established, yet published studies use at least 18 distinct calculation methods across 26 biomarkers, raising a fundamental question: does this association reflect a stable biological signal or an artifact of investigator choice? We applied the first multiverse specification-curve analysis of AL to two independent NHANES cohorts (NHANES III: n = 17,285; NHANES 2007-2010: n = 12,729) linked to the National Death Index through 2019, constructing 450 analytical specifications by systematically varying biomarker panel, scoring method, covariate configuration, and mortality outcome. Every cardiovascular mortality specification (100% of 150) and 93.3% of all-cause specifications produced statistically significant hazard ratios exceeding 1.0. Median HRs per 1-SD increment in AL were 1.22 (all-cause, pooled) and 1.36 (cardiovascular) -- numerically identical to Parker et al.s conventional meta-analytic estimate of 1.22. Biomarker panel composition explained 46% of between-specification variance, compared with 4% for covariate adjustment. A five-biomarker panel (SBP, BMI, HDL, HbA1c, creatinine) performed comparably to an 18-biomarker expanded panel, with 100% of specifications significant across both cohorts. The AL-mortality association is a robust biological signal; biomarker panel selection -- not covariate adjustment -- is the primary target for field-wide standardization.