Mechanism Matters: A Monte Carlo Evaluation of Estimator Validity and Collider Bias in Environmental Mixture Epidemiology
Obeng-Gyasi, E.
Show abstract
Background: Mixture epidemiology deploys sophisticated estimators, Bayesian kernel machine regression with causal mediation analysis (BKMR-CMA), quantile G-computation (QGC), and parametric G-computation, alongside conventional regression. Comparative evaluations have assumed additive, non-mediated data-generating processes, leaving conditions under which estimator choice determines causal validity uncharacterized. Methods: We developed a simulation framework using military-relevant exposure distributions (metals, per- and polyfluoroalkyl substances [PFAS], polychlorinated biphenyls [PCBs]) and allostatic load (AL) across three deployment tiers, with parameters drawn from military occupational health and contamination literature. Four data-generating processes were specified as directed acyclic graphs: direct effects with confounding (M1), full mediation through AL (M2), synergistic AL-exposure interaction (M3), and collider structure (M4). We evaluated ordinary least squares (OLS), QGC, G-computation, and BKMR-CMA on bias, root mean squared error, and 95% confidence interval coverage across 500 Monte Carlo replications at n = 500 and n = 1,000. Results: No estimator dominated across all mechanisms. Under M1, OLS and G-computation produced near-identical modest positive bias; BKMR-CMA achieved lower root mean squared error through kernel shrinkage. Under M2, BKMR-CMA exhibited severe positive bias for AL (mean bias = +0.579 SD units; coverage = 32.8%). Under M3, BKMR-CMA was the only estimator achieving nominal 95% coverage for AL (95.2%), while regression-based approaches fell to 83.6%. Under M4, G-computation produced persistent bias and near-zero coverage for lead, reflecting structural non-identification. Conclusions: Estimator validity is fundamentally mechanism-dependent. Researchers should base estimator choice on explicit causal assumptions about whether AL functions as confounder, mediator, moderator, or collider, particularly in military and occupational cohorts. We provide a mechanism-to-estimator mapping for applied researchers.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Analyses using multiple imputation need to consider missing data in auxiliary variables 93%
- Joint Effects of Indoor Air Pollution and Maternal Psychosocial Factors During Pregnancy on Trajectories of Early Childhood Psychopathology 92%
- Spatiotemporal Forecasting of Opioid-related Fatal Overdoses: Towards Best Practices for Modeling and Evaluation 91%
Similar papers in this journal
Similar papers in this journal
- Negative Control Exposures: Causal effect Identifiability and Use in Probabilistic-Bias and Bayesian Analyses with Unmeasured Confounders 95%
- Sensitivity and Uncertainty Analysis for Two-Stream Capture-Recapture Methods in Disease Surveillance 91%
- Assessing Direct and Spillover Effects of Intervention Packages in Network-Randomized Studies 89%
Similar papers in this journal
- Environmental Mixtures Analysis (E-MIX) Workflow and Methods Repository 95%
- Using parametric g-computation to estimate the effect of long-term exposure to air pollution on mortality risk and simulate the benefits of hypothetical policies: the Canadian Community Health Survey cohort (2005 to 2015) 94%
- Exposure contrasts of pregnant women during the Household Air Pollution Intervention Network randomized controlled trial 92%
Similar papers in this journal
- Efficient Estimation of Indirect Effects in Case-Control Studies Using a Unified Likelihood Framework 93%
- Mendelian Randomization with longitudinal exposure data: simulation study and real data application 93%
- Penalized reduced rank regression for multi-outcome survival data supports a common metabolic risk score for age-related diseases 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.