Back

Enhancing Bayesian Kernel Machine Regression: A Dynamic Thresholding Framework to Address Variability and Skewness in High-Dimensional Environmental Health Data

Hasan, K. T.; Odom, G.; Bursac, Z.; Lucchini, R.; Ibrahimou, B.

2025-04-16 occupational and environmental health
10.1101/2025.04.14.25325822 medRxiv
Show abstract

Bayesian Kernel Machine Regression (BKMR) is widely used in environmental health research to model complex, nonlinear, and interactive relationships in high-dimensional datasets. However, using a fixed posterior inclusion probability (PIP) threshold can lead to inconsistent test size control, influenced by the coefficient of variation (CV) and sample size. This study introduces a dynamic thresholding approach that adapts to these dataset characteristics, improving the sensitivity and reliability of BKMR analyses. A four-parameter logistic regression model was developed to estimate the 95th percentile of PIP as a function of the log-transformed CV and sample size. Simulations were performed across a broad range of CV values and sample sizes to evaluate test size performance for fixed and dynamic thresholds. The dynamic threshold was validated with independent simulated datasets and applied to the 2011-2014 NHANES data. The dynamic threshold consistently maintained nominal test sizes near five percent, outperforming the fixed threshold, which exhibited substantial variability. Validation with empirical data identified cadmium, manganese, and lead as significant contributors to cognitive performance, with cadmium emerging as the most influential. The dynamic threshold approach improves the precision of variable selection in BKMR, offering a more reliable method to analyze complex exposure-response relationships in environmental health research.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.