Back

Biostatistics

Oxford University Press (OUP)

Preprints posted in the last 90 days, ranked by how well they match Biostatistics's content profile, based on 24 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
MIDFA: Scalable Bayesian Factor Analysis for Mixed and Incomplete Data

Hutchings, G.; Samartsidis, P.; Donnay, C.; Gaetano, L.; Fisher, E.; Nichols, T. E.; Holmes, C.; Häring, D. A.; Ganjgahi, H.

2026-06-12 neuroscience 10.64898/2026.06.10.731315 medRxiv
Top 0.1%
39.4%
Show abstract

Probabilistic latent variable models are a powerful tool for uncovering structure in high-dimensional datasets, particularly in biomedical applications. The increasing availability of large-scale epidemiological studies, such as the UK Biobank, poses important modelling challenges, including mixed data types, high dimensionality, and structured missingness. Existing approaches address some of these issues, but few provide a unified and scalable framework for handling them simultaneously. Here, we propose a scalable Bayesian factor analysis framework designed to address these challenges. Our method combines a semi-parametric Gaussian copula model with a continuous spike-and-slab prior to induce sparse and interpretable factor loadings. The number of latent dimensions is learned nonparametrically from the data using an Indian buffet process prior. For model fitting, we develop an expectation-maximisation algorithm that naturally accommodates missing data. We validate the proposed method through comprehensive simulation studies. In addition, we showcase the proposed model using the Novartis-Oxford Multiple Sclerosis dataset in two ways. First, we identify latent dimensions shared across MS clinical and neuroimaging variables, characterising disease structure while demonstrating the models ability to handle multiple data types and structured missingness. Second, we use the model for dimensionality reduction of structural MRI data, extracting features for downstream analysis that go beyond traditional whole-brain summary statistics. In those applications, our method identifies sparse latent structures and provides insights beyond those obtained from traditional approaches.

2
Estimating Time-Inhomogeneous Transition Probabilities from District Level Daily COVID-19 Transmission Data in Sierra Leone

Jah, A.; Ngesa, O.; Wamwea, C.; Ngunyi, A.

2026-07-30 infectious diseases 10.64898/2026.07.28.26359161 medRxiv
Top 0.1%
29.9%
Show abstract

Abstract: Compartmental epidemic models conventionally treat the probability of moving between dis ease states as fixed over time, an assumption that sits uneasily with the reality of a pandemic in which lockdowns, mask mandates, vaccination roll-out, and the arrival of new variants continually reshape transmission. This paper develops a time-inhomogeneous Markov chain framework for the Susceptible-Exposed-Infectious-Removed (SEIR) process, in which each transition probability pab(t) is allowed to vary with calendar time while respecting the struc tural zeros implied by the SEIR compartmental flow. We derive the constrained maximum likelihood estimator of pab(t) under these structural constraints, establish its finite sample efficiency, asymptotic normality, and Wilson score confidence intervals, and construct a like lihood ratio test of the null hypothesis that a compartments exit probability is constant over time. We further propose a stochastic machine learning hybrid extension in which the raw, kernel smoothed transition probabilities are regressed on policy and mobility covariates using both a logistic generalized linear model and a random forest, allowing the framework to attribute time-inhomogeneity to observable interventions. The methodology is applied to a compiled daily, district level COVID-19 surveillance panel for Sierra Leone spanning March 2020 to December 2023 (16 districts, 1,401 days). The likelihood ratio test rejects time-homogeneity of the exposed to infectious transition in 15 of 16 districts and of the infectious-to-removed transition in 8 of 16 districts ( = 0.05), and the covariate augmented logistic model achieves an out of sample Brier score roughly 76 times smaller than a time homogeneous pooled baseline, with healthcare capacity and the time trend emerging as the most influential predictors in the random-forest component. These results provide statisti cal evidence that time-inhomogeneous, covariate informed Markov models offer a materially better description of district-level COVID-19 transmission in Sierra Leone than classical time homogeneous compartmental models, with implications for sub-national outbreak monitoring in resource constrained settings

3
An M-learner approach for heterogeneous mediation analysis with high-dimensional omics mediators

Li, X.; Wei, P.

2026-09-01 bioinformatics 10.64898/2026.08.25.747106 medRxiv
Top 0.1%
7.6%
Show abstract

Causal mediation analysis is widely used to identify biological pathways linking exposures to outcomes, but most methods assume homogeneous mediation effects across individuals. In high-dimensional omics settings, this assumption can mask important heterogeneity driven by demographic, genetic, or environmental factors. We propose the M-high-learner, a flexible framework for detecting heterogeneous mediation effects with high-dimensional mediators. The method identifies mediators with subgroup-specific indirect effects while distinguishing them from null or homogeneous signals and controlling the type I error rate. It is computationally efficient, scalable, and yields interpretable sub-types. Simulation studies show that the proposed approach achieves high power while maintaining accurate error control. Applications to the Framingham Heart Study and the Multi-Ethnic Study of Atherosclerosis reveal that the mediation role of gene expression in sexs effect on high-density lipoprotein varies across subgroups defined by body mass index and age. Our framework provides a practical tool for uncovering heterogeneous biological mechanisms in high-dimensional genomic studies. Author SummaryBiological processes linking risk factors to disease often differ across individuals, but many existing methods assume these processes are the same for everyone. This can hide important differences between groups. We developed a powerful method to identify when these pathways vary across subgroups using large-scale molecular data. Our approach detects differences in how intermediate biological factors contribute to outcomes in populations defined by characteristics such as age and body mass index. Applying our method to population studies, we found that some biological pathways operate differently across groups, suggesting that key mechanisms may be missed when differences are ignored. Our work provides a tool to better understand how disease-related processes vary across individuals, which may support more targeted and personalized approaches to health research.

4
SMS: Symmetric Mediation Statistics for Powerful High-Dimensional Mediation Analysis

Wang, Y.; Yan, S.; Wang, H.-J.; Hu, Y.-J.

2026-06-15 bioinformatics 10.64898/2026.06.10.730748 medRxiv
Top 0.1%
6.2%
Show abstract

BackgroundMediation analysis of high-dimensional features, particularly molecular-level omics features, provides important opportunities to uncover biological mechanisms underlying human health and disease. However, two central statistical challenges remain: testing the composite-null hypothesis and maintaining power when the exposure-mediator and mediator-outcome associations differ substantially in statistical significance. Existing methods typically rely on accurate estimation of the proportions of the three null types or on the maximum of the two association p-values, and may not always control the FDR well and may have limited power under imbalanced significance. MethodsWe propose SMS, a new statistical framework based on symmetric mediation statistics. By exploiting symmetry, SMS calibrates the composite null distribution as a whole for FDR control. It also allows flexible combinations of the two association p-values, including the maximum, and then enables construction of an omnibus test. Moreover, it permits direct use of effect-size estimates, bypassing the need to compute p-values. ResultsSMS controlled the FDR across a wide range of simulation scenarios while achieving a substantial sensitivity gain, often around 20 percentage points, over existing methods including HDMT, DACT, and DEI-B. Applications to a metabolomics dataset and a DNA methylation dataset further corroborated these findings. Notably, SMS discovered five plausible mediators in the metabolomics dataset that were missed by all existing methods considered.

5
Non-Parametric Ancestry Adjustment for Polygenic Scores

Mas Montserrat, D.; Barrabes, M.; Bustamante, C. D.; Ioannidis, A. G.

2026-06-15 genetic and genomic medicine 10.64898/2026.06.07.26355080 medRxiv
Top 0.1%
5.5%
Show abstract

Modern polygenic risk scores (PRS) exhibit shifts correlated with ancestry, leading to erroneous predictions for non-European individuals when models are trained on predominantly European cohorts. Such shifts arise from, among other factors, (1) algorithmic limitations in the ability of PRS model training to detect causal variants, rather than nearby variants with ancestry-dependent correlations to the causal one, (2) under-representation of alleles with higher prevalence in non-European populations in the association study training, and (3) gene-by-environment interactions where the environment is correlated with genetic ancestry. Current ancestry-adjustment methodologies often discretize individuals into population categories and apply a simple affine mapping to reduce these genetic ancestry biases. However, such approaches provide suboptimal adjustments, particularly for admixed individuals. In this work, we introduce a detailed theoretical characterization of ancestry-dependent biases and propose novel methods based on non-parametric neighborhood techniques that provide more accurate empirical results and admit statistical consistency guarantees. Extensive experiments using the UK Biobank demonstrate the effectiveness of the proposed methods.

6
GPSNorm: Gaussian Process Spatial Normalization for Spatial Transcriptomics

Taychameekiatchai, A.; Zhan, X.; Xiao, G.; Ruan, P.

2026-07-27 bioinformatics 10.64898/2026.07.23.740319 medRxiv
Top 0.1%
5.4%
Show abstract

Spatial transcriptomics technologies enable measurement of gene expression while preserving spatial tissue organization, but they remain highly sensitive to technical variability such as library size differences, slide-level effects, and spatial artifacts. Most existing normalization approaches treat normalization as a preprocessing step and perform downstream analyses on normalized values as fixed inputs, ignoring the uncertainty introduced during normalization. We introduce GPSNorm (Gaussian Process Spatial Normalization), a Bayesian spatial normalization framework that jointly models technical variation, spatial structure, and biological signal within a unified hierarchical model. GPSNorm represents gene expression counts using a negative binomial latent Gaussian model whose spatial component is a Gaussian Markov random field approximating a Gaussian process and performs efficient approximate Bayesian inference using the Integrated Nested Laplace Approximation (INLA), producing posterior estimates that propagate normalization uncertainty into downstream differential expression analysis. In simulations anchored to empirical spatial transcriptomics data, GPSNorm accurately recovers spatial technical structure and improves log-fold change estimation compared with existing normalization methods. Applications to three spatial transcriptomics datasets--including human dorsolateral prefrontal cortex Visium data, a GeoMx COVID-19 lung damage study, and the Spatial Organ Atlas kidney dataset--demonstrate improved preservation of biologically expected spatial patterns and marker gene contrasts. These results show that jointly modeling normalization and downstream inference can improve the robustness and interpretability of spatial transcriptomics analyses.An open-source R implementation of GPSNorm is available at https://github.com/Tiny-Quant/GPSNorm.

7
Prediction of fMRI activity using vector autoregressive models: a comparison of sparse and low-rank approaches

Tian, X.; Gibberd, A.; Roy, S.; Nunes, M.

2026-06-15 neuroscience 10.64898/2026.06.11.731556 medRxiv
Top 0.1%
5.3%
Show abstract

Vector autoregressive (VAR) models have a history of being used to examine functional connectivity in the brain, as captured by functional MRI studies. Such models allow for an estimation of Granger-causal relationships between regions of interest across the brain. Unfortunately, since the number of parameters in the VAR model scales as the square of the number of regions, and this is typically large compared to the number of temporal observations, these parameter estimates will exhibit high variance. To address this challenge, we introduce a low-rank pre-smoothing method that applies a low-rank approximation to the observations before fitting a VAR model. We estimate these models using individual subject data from both task-based and resting-state conditions, tuning hyperparameters at the population level. Our low-rank approach is directly compared against sparse and unconstrained estimation methods. Evaluations of predictive performance and model structure reveal that our pre-smoothing technique enables robust individual-level parameter estimation and significantly reduces prediction error, a finding further validated by synthetic experiments where the ground-truth parameters are known.

8
From exact gradients to exact standard errors: third-order sensitivity equations for the FOCE and FOCEI population likelihood

van de Beek, H.; Beldjenna, M.; Fidler, M. L.; Zwep, L. B.; van Hasselt, J. G. C.

2026-07-15 pharmacology and toxicology 10.64898/2026.07.09.737434 medRxiv
Top 0.1%
4.1%
Show abstract

Asymptotic standard errors for the parameters of a nonlinear mixed-effects model fitted by first-order conditional estimation (FOCE) or FOCE with interaction (FOCEI) require the observed (Fisher) information -- the negative second derivative of the population objective at the optimum. The gradient of this objective can be computed exactly from sensitivity equations, but the observed information is conventionally still formed by finite differencing, which is less accurate and step-size dependent. Our objectives are to (i) derive the FOCE and FOCEI observed information in closed form within the same sensitivity-equation framework, and (ii) quantify the precision this recovers. Writing the objective as a data term plus the log-determinant of the first-order inner Hessian, the population Hessian splits so that the data term reuses the second-order sensitivities already needed for the gradient, whereas the log-determinant term requires third-order sensitivity equations -- confining the third-order dependence to a single term, where it enters in exactly two places. A finite-difference error analysis shows the differenced Hessian attains an accuracy no better than the square root of the objectives evaluation accuracy, whereas the analytic form is limited only by the sensitivity and differential-equation solutions, with no step size to tune. We illustrate this on a one-compartment oral model with first-order absorption fitted to warfarin data, where the differenced standard error is usable only over a narrow band of step sizes while the analytic value carries none. Implemented in the open-source R package nlmixr2, the method makes exact, reproducible standard errors routine, supporting more dependable confidence intervals, identifiability assessment, and uncertainty propagation.

9
Bayesian Factor Analysis for Binary and Ordinal Phenotypes with Missingness

Shashaank, N.; Knowles, D. A.

2026-07-21 bioinformatics 10.64898/2026.07.15.733660 medRxiv
Top 0.1%
4.0%
Show abstract

Binary and ordinal phenotypes are common in clinical screening and self-reported questionnaires, but many factor analysis and matrix factorization methods are only applicable for quantitative phenotypes with real-valued and/or continuous data distributions. To address this, we propose FABOr (Factor Analysis of Binary and Ordinal data), a Bayesian framework for matrix factorization in which the low-rank matrices are latent variables with continuous priors while the phenotypes are observed variables modeled with appropriate binary/ordinal likelihoods. We also develop missing not at random (MNAR) extensions of FABOr for analyzing data with structured missingness. In experiments with simulated phenotypes, we found that FABOr performs similarly to the best-performing benchmark methods on binary data at imputation and exceeds the performance of all tested benchmark methods on ordinal data. We then applied FABOr to analyze real-world binary and ordinal phenotypes from the Simons Foundation SPARK dataset on autism spectrum disorder (ASD) and found that it improved imputation accuracy by up to 5% on binary data and up to 23% on ordinal data relative to the benchmark methods.

10
Competing event regression on the relative subdistribution and cumulative-incidence scales

Mell, L. K.

2026-08-14 epidemiology 10.64898/2026.08.13.26360204 medRxiv
Top 0.1%
4.0%
Show abstract

In competing risks settings, covariate effects and group comparisons are usually assessed one event at a time - through log-rank or Cox tests on the cause-specific hazards, or Gray's test or Fine-Gray regression on a cumulative incidence function (CIF). This can obscure a clinically important quantity: the ratio between the event of interest and the competing event, since groups may differ little on the individual events yet differ sharply in their ratio. The generalized competing event (GCE) framework makes this ratio the object of inference; on the cause-specific scale the hazard ratio omega+(t) = lambda_1(t)/lambda_2(t) is estimated efficiently from a single stacked (Lunn-McNeil) model. We extend the framework to two scales that describe realized incidence. The subdistribution hazard ratio omega-tilde+(t) = lambda-tilde_1(t)/lambda-tilde_2(t) is estimated by a stacked, risk-set-weighted extension of the Lunn-McNeil construction; the cumulative-incidence ratio rho(t) = F_1(t)/F_2(t) - the odds that a subject's realized event by time t is the event of interest - by jackknife pseudo-observation regression of the Aalen-Johansen estimator. We relate the three contrasts: rho equals omega+ exactly under proportional cause-specific hazards, and equals omega-tilde+ only in the small-time limit under proportional subdistribution hazards, drifting toward 1 thereafter. The orthogonality that makes omega+ efficient is lost on both cumulative-incidence scales - omega tilde+ through overlapping weighted risk sets and shared censoring weights, rho through the shared all-cause survivor - so each carries a covariance term that must be handled and that bounds efficiency relative to the hazard-scale test. We derive the corresponding variances, study operating characteristics by simulation, illustrate on hypothetical prostate and head-and-neck cohorts, and provide an implementation in the gcemod R package.

11
The Variance-Stabilizing Transformation for the Poisson Rate Ratio: Closed-Form Confidence Intervals

Ng, S.-P.

2026-07-18 epidemiology 10.64898/2026.07.16.26358255 medRxiv
Top 0.1%
3.9%
Show abstract

The incidence rate ratio R is the standard measure for comparing event rates in clinical trials and epidemiology. In vaccine trials, the vaccine efficacy is VE = 1 - R. When events are rare, the two arm counts are Poisson. The estimator of R is heteroskedastic: its sampling variance changes with the data. So no fixed-width interval covers correctly everywhere. The usual log-Wald interval is undefined at zero events and covers poorly at small counts. Early vaccine and drug-safety readouts fall in exactly this regime. We show that a single reparameterization collapses this bivariate problem to an effective one-parameter family with a quadratic variance function, whose variance-stabilizing transformation is 2 arcsinh(sqrt(R)). The reduction yields a closed-form confidence interval for R. Its two leading errors, a curvature bias and the variability of the estimated scale, each admit a closed-form correction with no tuning constants. In a Monte Carlo study of our seven arcsinh variants and five competitors, the +Curve+Stu variant covers within 0.002 of the nominal 0.95 for about 50 control and 5 treatment events. Its width is on par with the best competitor. It avoids the conservatism and zero-count breakdown of log-Wald and MOVER. For moderate counts, we recommend this interval; for sparser data, our Bar-Lev and Enis count-shift variant is more robust. The result is a ready-to-use, closed-form interval for the low-count regime. We illustrate it on early Covid-19 vaccine-efficacy readouts and provide reference implementations in R and Python.

12
Impact of Inaccurately Labeled Data on the Performance of Multi-label Classification for Disease Recognition

Schmiegel, S.; Marchi, H.; Roechter, M.-H.; Rudwaleit, M.; Fuchs, C.

2026-07-23 health informatics 10.64898/2026.07.22.26358665 medRxiv
Top 0.1%
3.4%
Show abstract

The process of medical diagnostics is challenging, especially since patients can simultaneously suffer from several diseases with similar, contradictory, or even opposing diagnoses. Statistical prediction can support physicians in this task; however, the quality of data used for predicition as well as the chosen statistical model can affect the reliability of data-driven decision support. Data quality can, in particular, be reduced by incomplete medical diagnoses, that is, the termination of the diagnostic process once a patient has tested positive for one disease that explains the symptoms. When interpreting missing diagnoses as negative, this leads to potentially false negative health data. Another source of low data quality lies in diagnoses being made through a principle of elimination, i.e., after several negative results, one opts for the seemingly last remaining possibility. This may lead to false positive health data. In our work, we investigate how such inaccurately labeled data affects the predictive ability of multi-label classification (MLC) for disease recognition. Unlike single-label classification (SLC), MLC allows the simultaneous assignment of multiple diseases to a patient and can therefore describe clinical conditions more holistically. To that end, we conduct a synthetic-data simulation study as well as a real-data case study on the example of chronic pain patients. In this regard, we compare MLC performance on accurately and inaccurately labeled data. We manipulate the data such that it corresponds to different diagnostic test sensitivities as well as to different examination sequences, thus paying special attention to resulting uncertainty within the process of medical diagnostics. Our results show that inaccurate labeling substantially decreases MLC prediction ability. Furthermore, low diagnostic test-sensitivity, the order of disease examination and covariate effects have a strong impact on MLC performance. These findings contribute to a better understanding of the interplay and impact of diagnostic procedures, data documentation and interpretation, and statistical modeling. This underlines the need for careful data collection as a basis for model development; special consideration should be given to the extensive examination of patients as well as the targeted collection of covariates. This is particularly crucial when models are transferred into everyday clinical practice.

13
Symptom-conditioned prediction of comorbidity in the hEDS-POTS-MCAS triad

Papas, K.; Banerji, A. I.

2026-07-28 rheumatology 10.64898/2026.07.27.26358921 medRxiv
Top 0.1%
3.3%
Show abstract

Background. Hypermobile Ehlers-Danlos syndrome (hEDS), postural orthostatic tachycardia syndrome (POTS) and mast cell activation syndrome (MCAS) are reported to co-occur frequently. It is unclear how strongly, whether symptom profiles sharpen prediction of a second diagnosis given a first, and whether the published literature can support such inference at all. Methods. We pooled 22 published cohorts-aggregate prevalence data, no primary human-subjects data-using Bayesian hierarchical random-effects models on the logit scale, and propagated the resulting posteriors through naive and tempered symptom updating. We introduce a feasibility screen derived from the Frechet-Hoeffding bounds that tests whether separately pooled marginals can describe a single population, and we characterise the identifiability of latent class structure under disease-selected sampling. Results. Directed comorbidity is strongly asymmetric: {pi}POTS|hEDS = 46.6% (95% CrI [32.5,61.5]) against {pi}hEDS|POTS = 12.1% ([3.5,38.5]), a near-fourfold gap, with {pi}MCAS|POTS lowest at 3.9% ([0.7,21.6]). Prediction intervals exceed credible intervals throughout, indicating substantial between-cohort heterogeneity. The feasibility screen finds 26 of 102 testable cells (25.5%) incompatible with any joint distribution; critically, 21 of these fail the upper Frechet bound and are invisible to the one-sided screen that is the natural first implementation. Among cells surviving the screen, symptom evidence is informative in four of six directions-P(hEDS | POTS,S) rises from 12.1% to 49.7% on a four-symptom panel under tempered updating, and P(POTS | MCAS,S) from 49.5% to 82.1%-but inert in both hEDS-cohort directions. A pathway-dispersion contrast excludes zero in two of six directions, in opposite signs and by margins of 0.1-0.2 percentage points, consistent with chance at this number of comparisons. We show latent class structure is not identified from disease-selected aggregate data, and that the single cohort reporting trivariate structure (N = 8) yields an exactly balanced table (OR = 1.00, 95% CI [0.063, 15.99]). Conclusions. The pooled directional probabilities are usable as clinical priors, with intervals wide enough to preclude precision. Symptom-conditioned prediction is supported in some directions but not those most often invoked clinically, and every estimate rests on cohorts dominated by self-reported ascertainment. The principal methodological contribution is the two-sided feasibility screen: applied here it shows that a quarter of the testable literature cannot describe one coherent population, and that a one-sided implementation understates this sixfold. Keywords: hypermobile Ehlers-Danlos syndrome; postural orthostatic tachycardia syndrome; mast cell activation syndrome; comorbidity; Bayesian meta-analysis; random-effects model; Frechet bounds; identifiability; latent class analysis; collider bias

14
Forecasting Trajectories of Physiological Mechanics with Sparse Clinical Data Using a Data Assimilation and Machine Learning Hybrid

Wang, Y.; Stroh, J. N.; Ghosh, D.; Sirlanci, M.; Hripcsak, G.; Bennett, T. D.; Albers, D.

2026-07-24 health informatics 10.64898/2026.07.22.26358695 medRxiv
Top 0.1%
3.1%
Show abstract

Clinical decisions for determining optimal patient-specific interventions are complicated prediction tasks that rely on health care professionals' understanding of physiological mechanisms and their dynamics. These decisions are challenged by (a) observational data sparsity and (b) patient heterogeneity. Here, we focus on estimating and forecasting specific physiological properties--that are not explicitly present in clinical observations--to provide additional features using only data available bedside at the time of decision-making. Mechanistic models of physiological system(s), e.g., physiological ordinary differential equation (ODE) models, provide pathways to compensate for data sparsity by synchronizing the model with observations of an individual patient using data assimilation (DA). However, DA used in a standard computational workflow to estimate constant model parameters from presently-known data is less effective at optimizing state forecasts of the model governed by physiological processes that evolve before new observations are available. Stated simply, we cannot forecast the future evolution of the model because we cannot forecast model parameters. To support next-generation clinical decision support, we develop a new DA and machine learning (ML) hybrid pipeline to estimate and forecast individual future physiological processes by forecasting ODE model parameters. This pipeline overcomes model and DA workflow limitations by stacking a DA-estimated posterior empirical distribution of physiological parameters with longitudinal ML forecasting models. We work within the context of glycemic management in an ICU using EHR data to construct and test a use case. We use synthetic data and real-world clinical data to validate the integrated pipeline and quantify uncertainties.

15
Model-based Detection of Spatial Disease Boundaries Using Amortized Bayesian Inference

Wu, K. L.; Banerjee, S.

2026-06-24 epidemiology 10.64898/2026.06.21.26356187 medRxiv
Top 0.1%
2.8%
Show abstract

Disease boundary analysis identifies abrupt changes in health outcomes across geographic boundaries, guiding targeted public health interventions and outbreak surveillance. Current implementations often adopt a Bayesian "wombling" approach and largely rely on Markov Chain Monte Carlo (MCMC) posterior sampling, presenting scalability issues for large-scale disease surveillance. We leverage amortized Bayesian inference (ABI) to accelerate the detection of spatial health disparities between neighboring US counties by embedding neural posterior estimation within a Bayesian areal wombling framework. Exploiting the computational efficiency of ABI, we further introduce the Residual Disparity Elimination Target, a metric for the required reduction in mortality or prevalence for a region to eliminate a significant disparity with its neighbor. We analyze tracheal, bronchus, and lung cancer mortality rates across mainland US counties and achieve results concordant with MCMC analysis while scaling areal wombling to hundreds of outcomes and translating disparity detection into interpretable policy objectives.

16
Simulation of synthetic health records for assessment of causal inference methods for vaccine efficacy

Velasco Pardo, V.; Daines, L.; Katikireddi, S. V.; Ritchie, L.; Robertson, C.; Simpson, C. R.; McCowan, C.; Swallow, B.

2026-07-19 infectious diseases 10.64898/2026.07.17.26358308 medRxiv
Top 0.1%
2.8%
Show abstract

Background During the COVID-19 pandemic, public health agencies used near real-time observational data to answer questions regarding vaccine effectiveness. However, traditional observational methods do not allow conclusions regarding counterfactual scenarios to be drawn from clinical data. Counterfactuals, which are outcomes that would have occurred under alternative interventions, can be used to formally assess the causal effects of public health interventions on health outcomes while accounting for the effects of confounding. Ideally individual patient data is used for the development of counterfactuals. Low-fidelity synthetic data may be useful for advancing methodological development where governance and privacy constraints prohibit access to sensitive personal data. Methods We simulated synthetic datasets based on the EAVE-II COVID-19 platform which has been limited to use for surveillance purposes. EAVE-II includes almost all resident people in Scotland registered with qualified general medical practitioners. Patient characteristics were simulated to reflect the known distribution of the Scottish population, accounting for dependencies between variables. Each synthetic dataset was encoded to different realistic scenarios for EAVEII 'ground truth' vaccine rollout and effectiveness results, explicitly stating the causal and confounding mechanisms, using a statistically sound method based on marginal structural models. Synthetic datasets of 100,000 individuals were then generated across five confounding scenarios and five severe outcome types. Results In scenarios with weak confounding, both unweighted and inverse probability of treatment weighted (IPTW) logistic regression recovered the true causal parameters. As confounding strength increased, only weighted models recovered the true mechanism. Conclusions Low-fidelity synthetic datasets simulated from EAVE-II data analysts to build and test causal inference pipelines, develop novel analysis pipelines, and train new researchers while awaiting access to real data. We showed how to generate synthetic datasets from a marginal structural model under different confounding scenarios.

17
Learning Minimal Gene Programs for Disease-Aligned Representations

Madduri, A.; Patel, C. J.

2026-07-22 bioinformatics 10.64898/2026.07.17.739212 medRxiv
Top 0.1%
2.7%
Show abstract

Identifying small, interpretable gene sets that robustly capture disease-associated variation in singlecell transcriptomic data remains a central challenge for biological interpretation and experimental followup. In practice, commonly used differential expression and sparsity-based approaches often produce large, unstable gene lists that fail to generalize across patients due to strong donor-specific confounding. We study sparse gene selection for reconstructing donor-robust, disease-aligned cellular trajectories in real single-cell RNA-seq datasets. We introduce Sparse Linear Manifold Control (SLMC), a practical workflow that defines a disease-aligned score after removing donor-associated variation and selects minimal gene programs whose expression reconstructs this score. We focus on diagnosing the structure of the resulting reconstruction objective and evaluating selection strategies under realistic health data conditions. Across five human single-cell datasets spanning oncology and neurodegeneration, we find that the reconstruction objective exhibits strong diminishing returns, explaining why simple greedy selection methods perform well in practice. Under strict donor-heldout evaluation, greedy methods consistently outperform LASSO at small gene budgets and achieve accurate reconstruction with as few as 25 genes. Together, these results highlight how careful objective design and empirical evaluation enable robust and interpretable gene selection for disease-aligned representation learning in single-cell health data. Data and Code AvailabilityAll code used for data preprocessing, model training, evaluation, and feature importance analyses is available at: https://github.com/AdiVM/SLMC_single-cell. The singlecell data used in this study are publicly available human transcriptomic datasets generated by prior studies and accessible through the Gene Expression Omnibus (GEO). Analyses were performed using renal cell carcinoma single-cell RNA-seq data (GEO accession: GSE314072) and Alzheimers disease single-nucleus RNA-seq data from human cortex (GEO accession: GSE138852), with additional publicly available datasets used for cross-context diagnostic evaluation (GEO accessions: GSM8652069, GSE308624, and GSE227734). All datasets contain de-identified human samples and are available under standard publicuse terms via GEO. Institutional Review Board (IRB)This study analyzes de-identified, publicly available human transcriptomic data obtained from previously published studies. No new data were collected, and no identifiable private information was accessed. In accordance with institutional policy, this work was determined to constitute non-human subjects research and did not require additional IRB approval.

18
POISE: Spectral Inference of Parent-of-Origin Effects in Unlabeled Genomic Data

Hwang, I.; Talbot, A.; Head, T.; Trevino, C.; Wingo, T. S.; Kotlar, A. V.

2026-06-10 genetics 10.64898/2026.06.10.731310 medRxiv
Top 0.1%
2.7%
Show abstract

MotivationParent of Origin Effects (POEs), where the effect of an an allele on a phenotype differs based on maternal or paternal inheritance implicated in growth, metabolism, and neurodevelopment. Traditional tests for POEs require family data to determine parental origins of transmitted alleles. Given that such studies are expensive and time consuming compared to genome-wide association studies (GWAS), tests that function absent inheritance information are highly desirable. We develop a method, based on community detection from machine learning, that infers POEs via a spectral decomposition, obtains confidence intervals via a non-parametric bootstrap, and safeguards against confounding by non POE sources of variation. We refer to our method as Parent of Origin Inference via Spectral Estimation (POISE). ResultsWe demonstrate that POISE is well-calibrated under both Gaussian and heavy-tailed noise in simulation studies, with improved robustness to true POEs compared to existing covariance-based tests. POISE provides per-trait effect estimates with bias-corrected bootstrap confidence intervals and incorporates an information-theoretic minimum detectable effect size that filters unreliable estimates, conferring robustness to covariance-deflating variance QTL. We then apply POISE to GWAS data from the UK Biobank using BMI, LDL cholesterol, and HDL cholesterol. POISE recovers established POE loci and identifies 134 additional variants at genes implicated in lipid metabolism, immune regulation, and growth. Availability and implementationThe code for this method in Python is available at https://github.com/bystrogenomics/POISE.

19
Two-Sample Instrumental Variables under Population Mismatch: A Transportability Framework with Bias Diagnostics

Qian, Y.; Song, Y.

2026-06-26 epidemiology 10.64898/2026.06.15.26355602 medRxiv
Top 0.1%
2.6%
Show abstract

Instrumental variable (IV) methods are widely used in health and social sciences to estimate causal treatment effects among compliers. In certain research settings, the instrument-treatment association (first stage) and the instrument-outcome association (reduced form) are each estimated from a different dataset. Two-Sample Instrumental Variables (TSIV), proposed by Angrist and Krueger (1992), addresses this by combining first-stage and reduced-form estimates from separate data sources into a single causal effect estimate. However, TSIV identification requires that instrument compliance behavior be consistent across the two samples, a condition that is rarely verified in practice. We show mathematically and empirically that when compliance differs between samples, the raw TSIV estimator does not converge to the true Local Average Treatment Effect (LATE) and instead attenuates toward a predictably biased limit proportional to the ratio of first-stage compliance rates between the two samples. To address this, we formalize a framework for estimating LATE with TSIV under two key assumptions: (1) Covariate Overlap, requiring that the two samples share sufficient common support in their covariate distributions, and (2) Compliance Transportability, requiring that compliance behavior is identical across populations after conditioning on observed covariates. We consider a setting in which a health policy instrument and outcomes are recorded in administrative claims while treatment and covariates are collected in a survey. We use a C-statistic derived from pooled covariates to detect population mismatch and an Inverse Probability Weighting (IPW) correction that reweights the first-stage sample to approximate the administrative covariate distribution. In Monte Carlo simulations across eight scenarios calibrated to a survey-Medicaid setting, IPW-TSIV reduces bias in estimating the LATE, achieving 88% reduction in the primary scenario, 82% under severe selection, and 79% when state-level expansion policy drives compliance heterogeneity. We further validate this framework using the Oregon Health Insurance Experiment, where partitioning the public-use lottery data (N = 24,646) into two non-overlapping samples with substantively meaningful compliance heterogeneity yields a verifiable benchmark against the true causal effect. IPW-TSIV reduces mean absolute bias by 71.6% relative to the oracle S2-specific LATE across 10 independent replications (C-statistic = 0.78), outperforms naive TSIV in all 10 splits, and reduces mean bias relative to the full-data LATE from +0.016 to +0.008. This framework provides applied researchers with actionable diagnostic thresholds to detect sample mismatch, validate transportability assumptions, and determine when structural TSIV estimation is reliable.

20
Covariance Nonstationarity is Evident in Spatial Transcriptomics and Provides a New Categorization of Spatially Varying Genes

Velidi, P.; Wei, Z.; Nathoo, F.

2026-08-18 bioinformatics 10.64898/2026.08.10.743911 medRxiv
Top 0.1%
2.6%
Show abstract

BackgroundGaussian process models underlie many spatial transcriptomics tools but typically assume stationary covariance. While typically ignored, non-stationarity of spatial covariance in gene expression may correspond to tissue heterogeneity or cell aggregates. ResultsAcross 13 Visium datasets, we use approximate Bayes factors from R-INLA to compare stationary and non-stationary Matern covariance functions. Evidence for covariance non-stationarity appears in 3% to 50% of genes across tissue samples. We further characterize the power and false discovery rate of the Bayesian analysis of non-stationarity. We find that gene sets associated with immune, cytokine, and other effector functions are enriched among genes favoring non-stationary spatial covariance. ConclusionsCovariance stationarity is not a benign technical simplification in spatial transcriptomics; it is frequently violated, the violation is biologically structured, and it changes the definition and classification of spatially varying genes.