Back

Biometrics

Oxford University Press (OUP)

All preprints, ranked by how well they match Biometrics's content profile, based on 23 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
zifalsnm: Zero-Inflated Bayesian factor analysis model with skew-normal priors for modeling microbiome data

Panchasara, S.; Jankowski, H.; McGregor, K.

2025-12-10 genetics 10.64898/2025.12.07.692834 medRxiv
Top 0.1%
45.6%
Show abstract

MotivationAdvancements in next-generation sequencing have transformed our understanding of host-microbe interactions, revealing links between microbial composition and chronic conditions such as obesity, diabetes, IBD, and others. However, the analysis of microbiome data is complex due to its unique statistical characteristics. One primary objective is to achieve effective dimension reduction to manage high dimensionality while simultaneously accounting for the datas compositional nature and zero inflation. Although existing probabilistic models provide frameworks for composition estimation, they are often based on the assumption that log-ratio-transformed compositions are normally distributed. This assumption is problematic, as it often fails to capture the significant skewness inherent in these transformed compositions. ResultsWe propose a new model called the Zero-Inflated Factor Analysis Logistic Skew-Normal Multinomial (ZIFA-LSNM) model : a comprehensive Bayesian hierarchical framework designed to address the statistical challenges of microbiome data. ZIFA-LSNM integrates a zero-inflation component to handle excess zeros, employs factor analysis for dimensionality reduction, and, critically, utilizes skew-normal priors on the latent factors to explicitly model data asymmetry. Posterior inference is performed using a scalable and efficient variational inference algorithm. Through simulation studies and real data analysis, the ZIFA-LSNM model has shown to demonstrate superior performance in parameter recovery and composition estimation compared to its Gaussian-based counterparts. Availability and Implementationzifalsnm is implemented in a freely available R package: https://github.com/SaurabhP-MS/zifalsnm.git Supplementary InformationSupplementary material is available with this article.

2
Identifying Effect Modification of Latent Population Characteristics on Risk Factors with a Sparse Varying Coefficient Regression

Wang, R.; Fang, L.; Wang, Y.; Jin, J.

2024-12-05 genetics 10.1101/2024.11.30.626101 medRxiv
Top 0.1%
45.6%
Show abstract

Leveraging observational data to understand the associations between risk factors and disease outcomes and conduct disease risk prediction is a common task in epidemiology. While traditional linear regression and other machine learning models have been extensively implemented for this task, the associations between risk factors and disease outcomes are typically deemed fixed. In many cases, however, such associations may vary by some underlying features of the individuals, which may involve certain subpopulation characteristics and environmental factors. While data for these latent features may not be available, the observed data on risk factors may have captured some proportion of the variation in these features. Thus extracting latent factors from risk factors and incorporating this effect modification into the model may better capture the underlying data structure and improve inference. We develop a novel regression model with some coefficients varying as functions of latent features extracted from the risk factors. We have demonstrated the superiority of our approach in various data settings via simulation studies. An application on a dataset for lung cancer patients from The Cancer Genome Atlas (TCGA) Program showed that our approach led to a 6% - 118% increase in (AUC-0.5) for distinguishing between different lung cancer stages compared to the classic lasso and elastic net regressions and identified interesting latent effect modifications associated with certain gene pathways.

3
Transfer Learning for Survival-based Clustering of Predictors with an Application to TP53 Mutation Annotation

Liu, X.; Yan, H.; Shi, H.; Montellier, E.; Chi, E. C.; Hainaut, P.; Wang, W.

2025-10-06 genetics 10.1101/2025.10.06.680732 medRxiv
Top 0.1%
45.2%
Show abstract

TP53 is the most frequently mutated gene in human cancers, and germline mutations in TP53 cause Li-Fraumeni syndrome (LFS), a hereditary predisposition to diverse cancers. Accurate annotation of TP53 mutations based on their survival effects is critical for informed LFS patient management. Motivated by this need, we develop a new approach for Survival-based Clustering of Predictors (SCP) by identifying homogeneous coefficients in Cox regression. We formulate this task as a fusionpenalized Cox regression problem and provide an efficient computational algorithm. A nonconvex distance-to-set penalty is adopted to facilitate parameter tuning and improve estimation accuracy. To overcome data limitations, we further develop TLSCP, a transfer learning extension that borrows coefficient ranking information from a source dataset under the assumption of similar ranking patterns between source and target. TL-SCP integrates ranking information through weighted rank averaging, allowing flexibility in accommodating cohort heterogeneity while maintaining model simplicity. Simulation studies demonstrate TL-SCPs superior performance over SCP in clustering recovery and coefficient estimation. In the application of TP53 mutation annotation where we utilize non-LFS germline TP53 mutation carriers as a source cohort for the target LFS cohort, TL-SCP identifies biologically meaningful TP53 mutation clusters and offers improved clinical interpretability compared to experiment-based annotations.

4
An Exponential Scale Mixture Model for Metatranscriptomics Data with Application to Inflammatory Bowel Disease

Kim, H.; Ma, L.

2026-05-15 genomics 10.64898/2026.05.15.725552 medRxiv
Top 0.1%
30.5%
Show abstract

Metatranscriptomic (MTX) sequencing enables profiling of gene expression across microbial communities, providing a framework for linking genetic potential with functional activity. However, standard pipelines report normalized abundances rather than raw counts, limiting the use of count-based RNA-seq methods, while Gaussian-based alternatives rely on transformations and assumptions that are often poorly suited to MTX data. We propose a new modeling framework for differential expression analysis of MTX data, built on a scale mixture of exponential distributions, that incorporates DNA abundance to adjust for genomic potential, accommodates subject-specific random effects, treats zeros as left-censored, and employs a mixture prior to handle extreme sparsity. Applied to the IBDMDB multi-omics cohort, differential expression results vary substantially across models, including among Gaussian approaches with different pseudocount choices. Our approach identifies a distinct subset of candidate genes not detected by existing Gaussian methods; these may provide useful leads toward a novel understanding of transcriptomic patterns associated with dysbiosis in inflammatory bowel disease. Estimated dysbiosis effect directions are consistent between our model and Gaussian-based approaches, while effect sizes from our model tend to be larger in absolute value.

5
Estimating Direct and Spillover Vaccine Effectiveness with Partial Interference under Test-Negative Design Sampling

Jiang, C.; Fang, F.; Talbot, D.; Schnitzer, M.

2025-02-25 infectious diseases 10.1101/2025.02.24.25322826 medRxiv
Top 0.1%
30.1%
Show abstract

The Test-Negative Design (TND), which involves recruiting care-seeking individuals who meet predefined clinical case criteria, offers valid statistical inference for Vaccine Effectiveness (VE) using data collected through passive surveillance, making it cost-efficient and timely. Infectious disease epidemiology often involves interference, where the treatment and/or outcome of one individual can affect the outcomes of others, rendering standard causal estimands ill-defined; ignoring such interference can bias VE evaluation and lead to ineffective vaccination policies. This article addresses the estimation of causal estimands for VE in the presence of partial interference using TND samples. Partial interference means that the vaccination of units within the same group/cluster may influence the outcomes of other members of the cluster. We define the population direct, spillover, total, and overall effects using the geometric risk ratio, which are identifiable under TND sampling. We investigate various stochastic policies for vaccine allocation in a counterfactual scenario, and identify policy-relevant VE causal estimands. We propose inverse-probability weighted (IPW) estimators for estimating the policy-relevant VE causal estimands with partial interference under the TND, and explore the statistical properties of these estimators.

6
Robust Inference of Individualized Treatment Effect in Mendelian Randomization

Liang, M.; Wu, R.; Xiao, F.; Li, X.

2026-05-12 genetics 10.64898/2026.05.08.723855 medRxiv
Top 0.1%
29.5%
Show abstract

Mendelian randomization (MR) is widely used to draw causal conclusions in the presence of unmeasured confounding, but most MR analyses focus on average treatment effects and rely on strong assumptions. For precision medicine, the primary target is instead the individualized treatment effect (ITE); yet in MR, such effects are not point-identified under core IV assumptions, and valid inference is particularly challenging. We therefore propose a robust partial identification inference framework for ITE under MR allowing multiple instruments. Under minimal causal assumptions, we derive a sharp inference procedure for the intersection bounds of ITE by adopting a multiplier bootstrap procedure with data-adaptive bootstrap distribution shifting and heterogeneous variance adjustment. In theory, we prove that the proposed method achieves nominal coverage and asymptotic sharpness. Further, we extend the procedure to tolerate possible invalid IVs under a minimal proportion rule assumption by aggregating over instrument subsets while preserving coverage. Simulation studies demonstrate that the proposed methods attain nominal coverage and substantially shorter intervals than existing procedures. We illustrate the framework using data from the Alzheimers Disease Neuroimaging Initiative to assess heterogeneous causal effects of TREM2 expression on Alzheimers disease risk across education-defined subgroups.

7
De-biased sparse canonical correlation for identifying cancer-related trans-regulated genes

Huey, N.; Dutta, D.; Laha, N.

2024-08-19 genetics 10.1101/2024.08.15.608166 medRxiv
Top 0.1%
27.4%
Show abstract

SO_SCPLOWUMMARYC_SCPLOWIn cancer multi-omic studies, identifying the effects of somatic copy number aberrations (CNA) on physically distal gene expressions (trans-associations) can potentially uncover genes critical for cancer pathogenesis. Sparse canonical correlation analysis (SCCA) has emerged as a promising method for identifying associations in high-dimensional settings, owing to its ability to aggregate weaker associations and its improved interpretability. Traditional SCCA lacks hypothesis testing capabilities, which are critical for controlling false discoveries. This limitation has recently been addressed through a bias correction technique that enables calibrated hypothesis testing. In this article, we leverage the theoretical advancements in de-biased SCCA to present a computationally efficient pipeline for multi-omics analysis. This pipeline identifies and tests associations between multi-omics data modalities in biomedical settings, such as the trans-effects of CNA on gene expression. We propose a detailed algorithm to choose the tuning parameters of de-biased SCCA. Applying this pipeline to data on estrogen receptor (ER)-associated CNAs and 10,756 gene expressions from 1,904 breast cancer patients in the METABRIC study, we identified 456 CNAs trans-associated with 256 genes. Among these, 5 genes were identified only through de-biased SCCA and not by the standard pairwise regression approach. Downstream analysis with the 256 genes revealed that these genes were overrepresented in pathways relevant to breast cancer.

8
BMDD: A Probabilistic Framework for Accurate Imputation of Zero-inflated Microbiome Sequencing Data

Zhou, H.; Chen, J.; Zhang, X.

2025-05-12 genomics 10.1101/2025.05.08.652808 medRxiv
Top 0.1%
26.8%
Show abstract

Microbiome sequencing data are inherently sparse and compositional, with excessive zeros arising from biological absence or insufficient sampling. These zeros pose significant challenges for downstream analyses, particularly those that require log-transformation. We introduce BMDD (BiModal Dirichlet Distribution), a novel probabilistic modeling framework for accurate imputation of microbiome sequencing data. Unlike existing imputation approaches that assume unimodal abundance, BMDD captures the bimodal abundance distribution of the taxa via a mixture of Dirichlet priors. It uses variational inference and a scalable expectation-maximization algorithm for efficient imputation. Through simulations and real microbiome datasets, we demonstrate that BMDD outperforms competing methods in reconstructing true abundances and improves the performance of differential abundance analysis. Through multiple posterior samples, BMDD enables robust inference by accounting for uncertainty in zero imputation. Our method offers a principled and computationally efficient solution for analyzing high-dimensional, zero-inflated microbiome sequencing data and is broadly applicable in microbial biomarker discovery and host-microbiome interaction studies. BMDD is available at: https://github.com/zhouhj1994/BMDD. Author SummaryUnderstanding the microbes living in and on our bodies--the microbiome--relies on analyzing complex sequencing data. However, these data often contain many zeros, either because a microbe is truly absent or simply missed due to insufficient sampling. These missing values make it hard to accurately analyze microbial patterns and identify important differences between groups, especially for methods that work on a log scale. To address this, we developed a new method called BMDD that uses a more realistic model to impute the zeros. Unlike existing tools that assume each microbe follows an unimodal abundance distribution, BMDD allows for microbes to follow a bimodal distribution, so they could behave differently in different conditions. It provides not just a single guess, but a range of possible values to better reflect the uncertainty. Our testing shows that BMDD more accurately recovers the true microbial profiles and improves the ability to detect meaningful differences between groups. This method can help researchers gain clearer insights into how the microbiome affects health and disease.

9
Personalized Risk Prediction for Cancer Survivors: A Bayesian Semi-parametric Recurrent Event Model with Competing Outcomes

Nguyen, N. H.; Shin, S. J.; Dodd-Eaton, E. B.; Ning, J.; Wang, W.

2023-03-01 genetics 10.1101/2023.02.28.530537 medRxiv
Top 0.1%
22.9%
Show abstract

Multiple primary cancers are increasingly more frequent due to improved survival of cancer patients. Characteristics of the first primary cancer largely impact the risk of developing subsequent primary cancers. Hence, model-based risk characterization of cancer survivors that captures patient-specific variables is needed for healthcare policy making. We propose a Bayesian semi-parametric framework, where the occurrence processes of the competing cancer types follow independent non-homogeneous Poisson processes and adjust for covariates including the type and age at diagnosis of the first primary. Applying this framework to a historically collected cohort with families presenting a highly enriched history of multiple primary tumors and diverse cancer types, we have derived a suite of age-to-onset penetrance curves for cancer survivors. This includes penetrance estimates for second primary lung cancer, potentially impactful to ongoing cancer screening decisions. Using Receiver Operating Characteristic (ROC) curves, we have validated the good predictive performance of our models in predicting second primary lung cancer, sarcoma, breast cancer, and all other cancers combined, with areas under the curves (AUCs) at 0.89, 0.91, 0.76 and 0.68, respectively. In conclusion, our framework provides covariate-adjusted quantitative risk assessment for cancer survivors, hence moving a step closer to personalized health management for this unique population.

10
Testing and Estimating Causal Treatment Effect Heterogeneity in Observational Studies via Revised Deep Semiparametric Regression:A Lung Transplant Case Study

Yuan, S.; Zou, F.; Zou, B.

2026-04-17 bioinformatics 10.64898/2026.04.13.718254 medRxiv
Top 0.1%
22.7%
Show abstract

Lung transplantation programs must decide when bilateral lung transplantation (BLT) offers meaningful functional benefit over single lung transplantation (SLT). Because donor and recipient characteristics jointly shape outcomes, the BLT-SLT contrast may differ across patients. However, analyzing observational registries poses a statistical challenge: apparent subgroup differences can be artifacts of complex confounding, while true heterogeneity can be missed or poorly quantified. Using a large national registry, we investigate whether the BLT effect varies across recipients and identify clinically relevant profiles of benefit using post-transplant lung function measured by forced expiratory volume in 1 second (FEV1). We develop deepHTL, a framework that tests for treatment effect heterogeneity and estimates how the BLT-SLT effect varies with patient features. In extensive simulations designed to resemble registry-like confounding, deepHTL controls false positives for detecting heterogeneity and yields more accurate individualized effect estimates than common machine learning methods. In the lung transplant cohort, we find strong evidence of heterogeneity in the BLT-SLT effect on FEV1: younger, lower risk recipients with better baseline status show the largest FEV1 gains from BLT, whereas older, higher risk candidates exhibit diminished marginal benefit. These findings provide statistically grounded guidance for patient selection and allocation of scarce donor organs.

11
A Bayesian Informative Shrinkage Approach for Large-scale Multiple Hypothesis Testing (BISHOT): with Applications in Differential Analysis of Omics Data

Su, Y.; Clark, M. E. J. Z.; Wang, C.

2025-09-16 genetics 10.1101/2025.09.11.675690 medRxiv
Top 0.1%
22.5%
Show abstract

A major goal of many omics studies is to identify differential features, e.g. differentially expressed genes, between experimental groups. When performing differential analysis for a given dataset, relevant information from another platform or species is often available. Incorporating such prior information can help identify features that show consistent differential patterns across platforms or species, which are more likely to reflect shared biological processes, and thereby enhance the robustness and generalizability of the findings. However, existing differential analysis methods typically analyze only the data from the current study and do not leverage prior knowledge about the magnitude or direction of changes from other platforms or species. We address this challenge, and the associated multiple testing problem, using a Bayesian framework that enables the incorporation of prior knowledge obtained from different platforms or species. We propose a new test statistic, Bayesian Credible Ratio (BCR), based on a heteroscedastic global local shrinkage prior, and a new multiple testing criterion, sign-adjusted FDR (SFDR), that emphasize information regarding the direction of the differentially features. We prove that BCR achieves the largest count of sign-based true positives among all legitimate SFDR-controlling methods. Simulation results offer numerical evidence of its advantage compared to an empirical Bayesian method. The approach is demonstrated through the analysis of RNAseq and single-cell RNAseq datasets.

12
Estimation of total mediation effect for a binary trait in a case-control study for high-dimensional omics mediators

Kang, Z.; Chen, L.; Wei, P.; Xu, Z.; Li, C.; Yang, T.

2025-02-02 genomics 10.1101/2025.01.28.635396 medRxiv
Top 0.1%
22.4%
Show abstract

Mediation analysis helps uncover how exposures impact outcomes through intermediate variables. Traditional mean-based total mediation effect measures may suffer from the cancellation of opposite component-wise effects, and existing methods often lack the power to capture weak effects in high-dimensional mediators. Additionally, most high-dimensional mediation analysis methods have focused on continuous outcomes, with limited attention to binary outcomes, particularly in case-control studies. To fill this gap, we propose an R2 total mediation effect measure within the liability framework that offers a clear and intuitive causal interpretation, provides additional insights beyond the mean-based measures, and is invariant to disease prevalence. We develop a cross-fitted, modified Haseman-Elston regression-based estimation procedure tailored for mediation analysis in case-control studies, which can also be applied to cohort studies. Our estimator remains consistent in the presence of non-mediators and weak effects, as demonstrated in extensive simulations. Theoretical justification for consistency is provided under mild conditions and without requiring exact mediator selection. In a case-control substudy of the Womens Health Initiative involving 2150 individuals, we found that many metabolites were mediators with weak effects in the path from BMI to coronary heart disease, and we estimated that 89% (95% CI: 57%-100%) of the BMI-explained variation in underlying CHD liability is mediated by the measured metabolomics. The proposed estimation procedure is implemented in the R package "r2MedCausal", available on GitHub.

13
Estimation of Mediation Effect for High-dimensional Omics Mediators with Application to the Framingham Heart Study

Yang, T.; Niu, J.; Chen, H.; Wei, P.

2019-09-19 genomics 10.1101/774877 medRxiv
Top 0.1%
22.3%
Show abstract

Environmental exposures can regulate intermediate molecular phenotypes, such as gene expression, by different mechanisms and thereby lead to various health outcomes. It is of significant scientific interest to unravel the role of potentially high-dimensional intermediate phenotypes in the relationship between environmental exposure and traits. Mediation analysis is an important tool for investigating such relationships. However, it has mainly focused on low-dimensional settings, and there is a lack of a good measure of the total mediation effect. Here, we extend an R-squared (Rsq) effect size measure, originally proposed in the single-mediator setting, to the moderate- and high-dimensional mediator settings in the mixed model framework. Based on extensive simulations, we compare our measure and estimation procedure with several frequently used mediation measures, including product, proportion, and ratio measures. Our Rsq measure has small bias and variance under the correctly specified model. To mitigate potential bias induced by non-mediators, we examine two variable selection procedures, i.e., iterative sure independence screening and false discovery rate control, to exclude the non-mediators. We evaluate the consistency of the proposed estimation procedures and introduce a resampling-based confidence interval. By applying the proposed estimation procedure, we find that more than half of the aging-related variations in systolic blood pressure can be explained by gene expression profiles in the Framingham Heart Study.

14
BayICE: A hierarchical Bayesian deconvolution model with stochastic search variable selection

Tai, A.-S.; Tseng, G.; Hsieh, W.-P.

2019-08-12 genomics 10.1101/732743 medRxiv
Top 0.1%
19.2%
Show abstract

Gene expression deconvolution is a powerful tool for exploring the microenvironment of complex tissues comprised of multiple cell groups using transcriptomic data. Characterizing cell activities for a particular condition has been regarded as a primary mission against diseases. For example, cancer immunology aims to clarify the role of the immune system in the progression and development of cancer through analyzing the immune cell components of tumors. To that end, many deconvolution methods have been proposed for inferring cell subpopulations within tissues. Nevertheless, two problems limit the practicality of current approaches. First, all approaches use external purified data to preselect cell type-specific genes that contribute to deconvolution. However, some types of cells cannot be found in purified profiles and the genes specifically over- or under-expressed in them cannot be identified. This is particularly a problem in cancer studies. Hence, a preselection strategy that is independent from deconvolution is inappropriate. The second problem is that existing approaches do not recover the expression profiles of unknown cells present in bulk tissues, which results in biased estimation of unknown cell proportions. Furthermore, it causes the shift-invariant property of deconvolution to fail, which then affects the estimation performance. To address these two problems, we propose a novel deconvolution approach, BayICE, which employs hierarchical Bayesian modeling with stochastic search variable selection. We develop a comprehensive Markov chain Monte Carlo procedure through Gibbs sampling to estimate cell proportions, gene expression profiles, and signature genes. Simulation and validation studies illustrate that BayICE outperforms existing deconvolution approaches in estimating cell proportions. Subsequently, we demonstrate an application of BayICE in the RNA sequencing of patients with non-small cell lung cancer. The model is implemented in the R package \"BayICE\" and the algorithm is available for download.

15
Methods for detection of clusters of observations with an outlying correlation coefficient value

Desmet, L.; Venet, D.; Trotta, L.; Burzykowski, T.; Buyse, M.

2020-10-14 health systems and quality improvement 10.1101/2020.10.12.20211128 medRxiv
Top 0.1%
19.1%
Show abstract

Multivariate datasets with a clustered structure are the natural framework for, e.g., multicentre clinical trials. We propose a number of methods aimed at detecting clusters with outlying correlation coefficients. While the methods can be used in a variety of settings, we focus mainly on their application to central statistical monitoring of clinical trials. In particular, we consider the issue of detecting centers (or other clusters of patients such as regions) with outlying correlation coefficients for bivariate data in a multicenter clinical trial. It appears that, in that context, the proposed methods perform well, as we show by using a simulation study and a number of real life datasets.

16
On Multiply Robust Mendelian Randomization (MR2) With Many Invalid Genetic Instruments

Sun, B.; Liu, Z.; Tchetgen Tchetgen, E.

2021-10-26 epidemiology 10.1101/2021.10.21.21265317 medRxiv
Top 0.1%
18.9%
Show abstract

Mendelian randomization (MR) is a popular instrumental variable (IV) approach, in which genetic markers are used as IVs. In order to improve efficiency, multiple markers are routinely used in MR analyses, leading to concerns about bias due to possible violation of IV exclusion restriction of no direct effect of any IV on the out-come other than through the exposure in view. To address this concern, we introduce a new class of Multiply Robust MR (MR2) estimators that are guaranteed to remain consistent for the causal effect of interest provided that at least one genetic marker is a valid IV without necessarily knowing which IVs are invalid. We show that the proposed MR2 estimators are a special case of a more general class of estimators that remain consistent provided that a set of at least k{dagger} out of K candidate instrumental variables are valid, for k{dagger}[≤] K set by the analyst ex ante, without necessarily knowing which IVs are invalid. We provide formal semiparametric theory supporting our results, and characterize the semiparametric efficiency bound for the exposure causal effect which cannot be improved upon by any regular estimator with our favorable robustness property. We conduct extensive simulation studies and apply our methods to a large-scale analysis of UK Biobank data, demonstrating the superior empirical performance of MR2 compared to competing MR methods.

17
Individual-Specific Gaussian Graphical Models for Heterogeneous Populations with Application to Epigenetic Gene Regulation in Lung Adenocarcinoma

Saha, E.

2026-05-28 bioinformatics 10.64898/2026.05.25.727641 medRxiv
Top 0.1%
18.7%
Show abstract

Inter-patient molecular heterogeneity is a fundamental challenge in precision oncology: population-level multi-omics networks reveal average biology aggregated across the population but obscure individual variations that drive differential clinical outcomes. We introduce SIREN (Sample-specific Inference via Regularized Empirical-Bayes Networks), a method that estimates one partial correlation network per sample across omics layers by combining a population-level empirical Bayes prior with a rank-1 individual-specific update. Since a sample-specific precision matrix cannot be estimated from a single observation, SIREN uses a conjugate Inverse Wishart prior whose mean is the Oracle Approximating Shrinkage estimator, yielding closed-form individual-specific posteriors without MCMC. On simulated heterogeneous populations, SIREN achieves superior edge recovery over population-average methods including OAS, Ledoit-Wolf, and graphical Lasso, while remaining competitive in homogeneous settings. Applied to paired transcriptomic and methylomic profiles from lung adenocarcinoma, SIREN identifies individual-specific gene-methylation regulatory edges that stratify patients by survival in ways population-level analysis cannot, implicating chromatin remodeling and WNT signaling pathways in epigenetic heterogeneity. SIREN is computationally scalable and available as a Python package.

18
Model-based Detection of Spatial Disease Boundaries Using Amortized Bayesian Inference

Wu, K. L.; Banerjee, S.

2026-06-24 epidemiology 10.64898/2026.06.21.26356187 medRxiv
Top 0.1%
18.5%
Show abstract

Disease boundary analysis identifies abrupt changes in health outcomes across geographic boundaries, guiding targeted public health interventions and outbreak surveillance. Current implementations often adopt a Bayesian "wombling" approach and largely rely on Markov Chain Monte Carlo (MCMC) posterior sampling, presenting scalability issues for large-scale disease surveillance. We leverage amortized Bayesian inference (ABI) to accelerate the detection of spatial health disparities between neighboring US counties by embedding neural posterior estimation within a Bayesian areal wombling framework. Exploiting the computational efficiency of ABI, we further introduce the Residual Disparity Elimination Target, a metric for the required reduction in mortality or prevalence for a region to eliminate a significant disparity with its neighbor. We analyze tracheal, bronchus, and lung cancer mortality rates across mainland US counties and achieve results concordant with MCMC analysis while scaling areal wombling to hundreds of outcomes and translating disparity detection into interpretable policy objectives.

19
Privacy-Enhancing Sequential Learning under Heterogeneous Selection Bias in Multi-Site EHR Data

Kundu, R.; Shi, X.; Patel, K. K.; Ohno-Machado, L.; Salvatore, M.; Song, P. X. K.; Mukherjee, B.

2025-09-28 epidemiology 10.1101/2025.09.26.25336642 medRxiv
Top 0.1%
18.5%
Show abstract

ObjectiveTo develop privacy-enhancing statistical methods for estimation of binary disease risk model association parameters across multiple electronic health record (EHR) sites with heterogeneous selection mechanisms, without sharing raw individual-level data. We illustrate their utility through a cross-biobank analysis of smoking and 97 cancer subtypes using data from the NIH All of Us (AOU) and the Michigan Genomics Initiative (MGI). Materials and MethodsLarge-scale biobanks often follow heterogeneous recruitment strategies and store data in separate cloud-based platforms, making centralized algorithms infeasible. To address this, we propose two decentralized sequential estimators namely, Sequential Pseudo-likelihood (SPL) and Sequential Augmented Inverse Probability Weighting (SAIPW) that leverage external population-level information to adjust for selection bias, with valid variance estimation. SAIPW additionally protects against misspecification of the selection model using flexible machine learning based auxiliary outcome models. We compare SPL and SAIPW with the existing Sequential Unweighted (SUW) estimator and with centralized and meta learning extensions of IPW and AIPW in simulations under both correctly specified and misspecified selection mechanisms. We apply the methods to harmonized data from MGI (n = 50,935) and AOU (n = 241,563) to estimate smoking-cancer associations. ResultsIn simulations, SUW exhibited substantial bias and poor coverage. SPL and SAIPW yielded unbiased estimates with valid coverage probabilities under correct model specification, with SAIPW remaining robust under selection model misspecification. Both approaches showed no notable efficiency loss relative to centralized methods. Meta-learning methods were efficient for large sites but failed in settings with small cohort sizes and rare outcome prevalence. In real-data analysis, strong associations were consistently identified between smoking and cancers of the lung, bladder, and larynx, aligning with established epidemiological evidence. ConclusionOur framework enables valid, privacy-enhancing inference across EHR cohorts with heterogeneous selection, supporting scalable, decentralized research using real-world data.

20
Highly adaptive LASSO: Machine learning that provides valid nonparametric inference in realistic models

Butzin-Dozier, Z.; Qiu, S.; Hubbard, A. E.; Shi, J.; van der Laan, M.

2024-10-19 epidemiology 10.1101/2024.10.18.24315778 medRxiv
Top 0.1%
18.2%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWUnderstanding treatment effects on health-related outcomes using real-world data requires defining a causal parameter and imposing relevant identification assumptions to translate it into a statistical estimand. Semiparametric methods, like the targeted maximum likelihood estimator (TMLE), have been developed to construct asymptotically linear estimators of these parameters. To further establish the asymptotic efficiency of these estimators, two conditions must be met: 1) the relevant components of the data likelihood must fall within a Donsker class, and 2) the estimates of nuisance parameters must converge to their true values at a rate faster than n-1/4. The Highly Adaptive LASSO (HAL) satisfies these criteria by acting as an empirical risk minimizer within a class of cadlag functions with a bounded sectional variation norm, which is known to be Donsker. HAL achieves the desired rate of convergence, thereby guaranteeing the estimators asymptotic efficiency. The function class over which HAL minimizes its risk is flexible enough to capture realistic functions while maintaining the conditions for establishing efficiency. Additionally, HAL enables robust inference for non-pathwise differentiable parameters, such as the conditional average treatment effect (CATE) and causal dose-response curve, which are important in precision health. While these parameters are often considered in machine learning literature, these applications typically lack proper statistical inference. HAL addresses this gap by providing reliable statistical uncertainty quantification that is essential for informed decision-making in health research.