Back

Biostatistics

Oxford University Press (OUP)

Preprints posted in the last 30 days, ranked by how well they match Biostatistics's content profile, based on 24 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
Recalibrating Mendelian randomization under winner's curse, sample structure and polygenicity

Yang, Y.; Lin, Z.; Xue, H.; Zhu, X.

2026-07-07 genetic and genomic medicine 10.64898/2026.06.25.26356593 medRxiv
Top 0.1%
6.2%
Show abstract

Recently, Hu et al. (2024) conducted a benchmarking study showing that most existing Mendelian randomization (MR) methods exhibit substantial bias and inflated type-I error rates in real data. They attributed these failures to two largely neglected sources of bias: winner's curse and polygenicity-induced bias. Although a few methods have been developed to address one or both of these issues, existing approaches either do not fully account for both biases or are restricted to the univariable setting. In this paper, we propose a multivariable Rao-Blackwellization that corrects winner's curse while accounting for polygenicity and sample structure in a unified framework. Unlike univariable Rao-Blackwellization, where instrument selection yields a truncated normal statistic amenable to a Mills-ratio correction, multivariable Rao-Blackwellization conditions on a noncentral $\chi^2$ statistic, for which no analogous correction is available. We derive closed-form conditional moments under this instrument selection model and use them to construct bias-corrected summary statistics that can be integrated into a wide range of existing MR methods. Simulations and real data analyses show that, when combined with methods such as MR-cML and MR-BEE, the proposed correction substantially improves type-I error control and yields more robust inference.

2
Calibrating machine learning approaches for probability estimation without calibration data

Di Carluccio, E.; Koliopanos, G.; Ojeda, F. M.; Weimar, C.; Ziegler, A.

2026-07-13 epidemiology 10.64898/2026.07.10.26357723 medRxiv
Top 0.1%
4.7%
Show abstract

Statistical prediction models for binary outcomes are becoming increasingly popular. One significant challenge is calibrating these models to suit the characteristics of a target population that is structurally different from the original population. Calibration is especially challenging when there is no training data available from the target population. To address this problem, we propose a novel calibration method, SimCal, which uses synthetic data generated from the model development data in conjunction with marginal statistics from the calibration cohort. We show that expert judgment modeling (EJM) may be used for calibration if cross-sectional data from the target population are available comprising expert judgments about the potential outcome and the covariates. We describe three alternative calibration approaches when calibration data are lacking: similarity-binning averaging (SBA), adaptive calibration of predictions (ACP), and Elkan calibration. In a simulation study, we compare SBA, ACP, Elkan calibration, and SimCal. R code for applying these methods is provided from the re-analysis of data on coronary artery disease. We illustrate all 5 calibration approaches with a real data set for predicting functional outcome after stroke and all approaches but EJM in the re-analysis of the Cleveland Clinic data. None of the approaches performed convincingly well in all situations. SimCal performed well when model parameters were correctly specified. EJM failed on the stroke data. Further research is urgently required for calibration in the absence of calibration data.

3
Why linkage disequilibrium measures disagree: Fisher geometry of rare common haplotype structure

Ichikawa, Y.

2026-07-07 genetics 10.64898/2026.07.02.736022 medRxiv
Top 0.1%
4.3%
Show abstract

Conventional LD measures such as r2 perform poorly in the rare common regime, particularly in asymmetric configurations such as nested haplotype structure. Because r2 is symmetric and quadratic, it removes directional structure in two ways: squaring discards the sign, or phase, retained by the signed LD coefficient D, while symmetric normalization hides the asymmetry between the conditional probabilities P(A|B) and P(B|A). Although D recovers the phase, it is locus symmetric and unnormalized; its magnitude is hard to compare across frequency regimes and it does not by itself express which way the asymmetry runs. We therefore analyze the conditional-probability asymmetry {Delta} = P(A|B) - P(B|A), together with r2 and D, as distinct scalar functions on the haplotype simplex under the Fisher information metric. The conditional probabilities P(A|B) and P(B|A) are bounded in [0, 1], directly express carrier-set inclusion, and are more readily visualized than D. Moreover, their difference admits the exact decomposition {Delta} = M + C into a marginal frequency term M and an LD-coupled term C. Prior work has characterized either the mathematical behavior of LD normalizations across allele-frequency space or the Fisher geometry of the haplotype simplex, but not their connection. We bridge this gap by showing that the geometric structure of the simplex explains why LD measures disagree in the rare common regime and why symmetric normalizations such as r2 lose directional information. We show that the fixed-frequency leaf is intrinsically anisotropic, positively curved, and frequency-dependent under the Fisher metric. These geometric predictions are tested empirically , in phased 1000 Genomes data1 and a two locus Wright Fisher model, in a companion paper (Ichikawa, preprint); the present note develops the geometry itself. Keywords: linkage disequilibrium; Fisher information metric; haplotype simplex; rare variant; conditional-probability asymmetry; nested haplotype structure

4
The Variance-Stabilizing Transformation for the Poisson Rate Ratio: Closed-Form Confidence Intervals

Ng, S.-P.

2026-07-18 epidemiology 10.64898/2026.07.16.26358255 medRxiv
Top 0.1%
3.2%
Show abstract

The incidence rate ratio R is the standard measure for comparing event rates in clinical trials and epidemiology. In vaccine trials, the vaccine efficacy is VE = 1 - R. When events are rare, the two arm counts are Poisson. The estimator of R is heteroskedastic: its sampling variance changes with the data. So no fixed-width interval covers correctly everywhere. The usual log-Wald interval is undefined at zero events and covers poorly at small counts. Early vaccine and drug-safety readouts fall in exactly this regime. We show that a single reparameterization collapses this bivariate problem to an effective one-parameter family with a quadratic variance function, whose variance-stabilizing transformation is 2 arcsinh(sqrt(R)). The reduction yields a closed-form confidence interval for R. Its two leading errors, a curvature bias and the variability of the estimated scale, each admit a closed-form correction with no tuning constants. In a Monte Carlo study of our seven arcsinh variants and five competitors, the +Curve+Stu variant covers within 0.002 of the nominal 0.95 for about 50 control and 5 treatment events. Its width is on par with the best competitor. It avoids the conservatism and zero-count breakdown of log-Wald and MOVER. For moderate counts, we recommend this interval; for sparser data, our Bar-Lev and Enis count-shift variant is more robust. The result is a ready-to-use, closed-form interval for the low-count regime. We illustrate it on early Covid-19 vaccine-efficacy readouts and provide reference implementations in R and Python.

5
netPCF: Geometry-Aware Pair Correlation Functions for Spatial Biology

Moore, J. W.; Bull, J. A.; Byrne, H. M.

2026-07-07 bioinformatics 10.64898/2026.07.02.736020 medRxiv
Top 0.2%
1.4%
Show abstract

Spatial organisation is a defining feature of biological systems, underpinning cellular interactions, tissue function, disease progression and therapeutic response. Identifying and quantifying spatial organisation may require methods that resolve relationships across spatial scales. The pair correlation function (PCF) quantifies spatial dependence between points across multiple length scales, but its standard Euclidean formulation is poorly suited to data defined on irregular, curved or otherwise structured domains, where tissue geometry may constrain biological organisation and distort Euclidean distances. Here, we introduce netPCF, a geometry-aware extension of the PCF for quantifying spatial organisation on complex biological domains. By representing tissue structures, anatomical surfaces and other constrained geometries as spatial networks, netPCF generalises the PCF beyond extrinsic Euclidean settings. The framework derives the expected behaviour of the statistic under complete spatial randomness using interpretable finite-support kernels, provides bootstrap-based uncertainty quantification, and includes practical criteria for assessing domain discretisation adequacy. We further extend netPCF to marked (labelled) biological data using feature kernels for categorical and continuous attributes, enabling unified analysis of cell identities, marker intensities, phenotypic states, gene expression and other quantitative features on structured domains in any spatial dimension. All methods are implemented in the open-source Python package spacenet. Synthetic studies show that netPCF recovers classical Euclidean behaviour on sufficiently resolved networks and is robust to common imaging noise. We demonstrate its utility in two biological applications. In three-dimensional imaging mass cytometry data from HER2+ breast carcinoma, netPCF separates tissue architecture-driven proximity from biologically meaningful endothelial and immune cell organisation. In reconstructed surfaces of developing murine embryos, netPCF identifies a transition in the Wnt1-Wnt6 relationship from short-range co-localisation at E9.5 to spatial exclusion at E11.5, a pattern of ectodermal boundary refinement not captured by prior voxel-wise co-expression analysis. Overall, netPCF provides a statistically grounded and practical framework for quantifying spatial organisation on complex biological domains.

6
Scalable multi-group nonnegative spatial factorization for spatial genomics data with cell-type heterogeneity

Chumpitaz-Diaz, L.; Shrestha, P.; Engelhardt, B. E.

2026-07-03 genomics 10.64898/2026.06.29.735224 medRxiv
Top 0.2%
1.3%
Show abstract

Spatial transcriptomics (ST) technologies enable the study of gene expression within the spatial context of tissues, providing insights into tissue structure, cellular interactions, and disease progression. However, existing dimension reduction methods often overlook spatial information or struggle to distinguish spatial gene patterns from those driven by cell-type differences, limiting biological interpretability by convolving differences in gene expression patterns with differences in cell-type proportions. To address these challenges, we introduce the scalable multi-group nonnegative spatial factorization (smNSF), a computationally-tractable probabilistic framework that integrates spatial coordinates and cell-type labels into a unified matrix factorization model. By using multi-group Gaussian processes (MGGPs) as priors, our model captures complex spatial variation in a cell-type specific way while enforcing nonnegativity to enhance interpretability. We develop a variational inference framework for MGGPs that supports scalable optimization and improves the numerical stability of smNSF. Across seven spatial transcriptomics datasets spanning diverse technologies and tissues, smNSF recovers sparse, interpretable spatial factors and, through its cell-type conditional posteriors, organizes them into cell-type enriched, cell-type specific, and universal spatial programs that are not apparent from marginal factors alone. Given cell-type labels in ST data, smNSF enables cell-type aware spatial decompositions and supports cell-type conditional posteriors for in silico exploration of relationships between spatial patterns and cellular identity.

7
Demographic Calibration Gaps in Breast Cancer Risk Prediction: Introducing the Demographic Calibration Gap Score

Eniolade, M.

2026-06-22 health informatics 10.64898/2026.06.17.26355900 medRxiv
Top 0.2%
1.1%
Show abstract

ABSTRACT: Most breast cancer prediction studies skip calibration reporting entirely. Fewer still examine calibration by demographic subgroup. Predicted probabilities that are systematically off for specific racial or gender groups produce biased clinical decisions, and aggregate statistics will not catch that. Objective: To introduce the Demographic Calibration Gap Score (DCGS), a metric that measures how much calibration error varies across demographic subgroups, and to show how it performs across five classifiers, four calibration conditions, and two datasets. Methods: Five classifiers were trained on the Wisconsin Diagnostic Breast Cancer dataset (n=569) and evaluated on a breast cancer cohort from MIMIC-IV (n=1,316). Three global calibration methods were applied: no calibration, Platt scaling, and isotonic regression. A fourth condition, subgroup-targeted Platt scaling, was applied to the MIMIC cohort. DCGS was computed as across racial and gender subgroups, with 95% bootstrap confidence intervals. Conformal prediction coverage and Demographic Coverage Gap (DCG) were reported. Results: On Wisconsin, all five models achieved AUROC above 0.98 and ECE below 0.12. Performance fell sharply on the MIMIC external cohort: AUROC dropped to 0.45-0.57 for base and globally calibrated variants, confirming distributional shift. DCGS exceeded the 0.05 clinical significance threshold in 28 of 40 model-calibration combinations on the race axis. Neither global Platt nor isotonic calibration reliably reduced DCGS below that threshold. Conformal coverage collapsed to roughly 25% on MIMIC, and racial DCG exceeded 0.15 for all 20 model-variant combinations. Conclusions: Reducing population-level ECE through global recalibration does not reliably close demographic calibration gaps. DCGS gives researchers a direct, standardized way to detect and report those disparities. Code and the DCGS computation library are released as open-source Python under the MIT License.

8
Overinflation and overconcentration: why Cauchy perturbation kernels are the right choice for ABC-SMC

Sturrock, M.; Shahrezaei, V.

2026-07-09 systems biology 10.64898/2026.06.24.734205 medRxiv
Top 0.2%
1.1%
Show abstract

Approximate Bayesian computation sequential Monte Carlo (ABC-SMC) propagates its particles with a perturbation kernel, and with the standard Normal kernel it degrades sharply as the parameter dimension grows, a failure usually attributed to dimension itself. We show instead that it is governed by the quality of the summary statistics, with dimension entering only through a separate and milder mechanism, and that the two must act together for the Normal kernel to break. The first ingredient is covariance overinflation: the kernel covariance, estimated from the particle cloud, overshoots the true posterior covariance by a factor set by information loss in the summary statistics. We derive this overscaling factor in closed form for a Gaussian model with sufficient statistics and show that it stays modest at any dimension, shrinking toward its baseline value as the tolerance tightens; the extreme values seen in practice (of order 103) are a signature of insufficient summaries, not of dimension. The second ingredient is perturbation overconcentration: the normalised Normal step size concentrates around one as the dimension grows, so every proposal overshoots by the same factor. Either ingredient alone is harmless; only their combination breaks the Normal kernel. A Cauchy kernel (multivariate t with one degree of freedom) removes the concentration, keeping a positive acceptance rate under arbitrary overscaling at a bounded worst-case cost of 1.87x in expected squared jump distance. In a Metropolis-Hastings framework we derive closed-form acceptance rates for both kernels that illustrate the advantage of the Cauchy kernel in this limit. A series of full ABC-SMC computational experiments on five problems at d = 12, including a hierarchical gene-expression model, show the Cauchy reducing the sliced Wasserstein distance to the reference posterior by factors of up to 50 with the same simulation budget. Since the summary statistics are commonly insufficient for the models that require ABC, overinflation is structural and the Cauchy perturbation kernel is the right default for problems in higher dimensions.

9
When Less Is Not More: DICEPro Mitigates the Impact of Incomplete Reference Matrices on Cellular Frequency Deconvolution.

BA, K.; Thiebaut, R.; Hinaut, X.; Hejblum, B. P.

2026-06-22 bioinformatics 10.64898/2026.06.17.732876 medRxiv
Top 0.2%
1.1%
Show abstract

Cellular deconvolution aims to estimate the frequencies of different cell populations from gene expression measurements in a biological sample. Supervised approaches, such as CIBERSORTx and DISSECT, critically depend on the reference signature matrix, which encodes the gene expression profiles of cell-types based on prior knowledge. Despite numerous deconvolution methods, the impact of missing cell populations in the reference matrix remains understudied. Here, we evaluate the robustness of state-of-the-art deconvolution approaches using simulations based on real dataset examples combined with statistical modeling, validated against published data, and multiple real benchmark datasets. Results show that deconvolution performance remains stable when the reference matrix includes most cell-types, but declines sharply as the matrix becomes incomplete, especially for abundant cell populations. To address the limitations of incomplete reference matrices, we introduce DICEPro, an optimization-based framework designed to enhance existing deconvolution methods. By systematically adjusting the reference signatures, DICEPro better accounts for missing or underrepresented cell populations, leading to improved precision and robustness. We show that DICEPro consistently boosts deconvolution performance across both simulated datasets, derived from real data examples, and multiple real biological datasets, offering a practical solution when standard methods are hindered by incomplete references.

10
Model-based inference of gene expression noise from single-cell RNA-sequencing data

Giersdorf, F.; Rogers, D. W.; Christensen, S.; Dutheil, J. Y.

2026-06-23 bioinformatics 10.64898/2026.06.18.733122 medRxiv
Top 0.2%
1.1%
Show abstract

The heterogeneity of expression levels among genetically identical cells, termed gene expression noise, is a property of the gene expression process whose importance in the biology of organisms and their evolution is increasingly recognized. Measuring gene expression noise requires single-cell expression data, as obtained from single-cell RNA sequencing (scRNASeq). Its estimation, however, is challenging owing to (i) the presence of technical noise in addition to biological noise, and (ii) the heterogeneity of cell types in the sampled population. We propose a maximum-likelihood framework to infer biological noise from scRNASeq data, while accounting for technical noise, dropout probabilities, and distinct cell sequencing depths. We demonstrate the parameter identifiability using simulations and that the resulting noise estimates are uncorrelated from the mean gene expression, and therefore do not need extra correction in downstream analyses, easing intra- and inter-genome comparisons. Using two technical replicates of scR-NASeq data from the wild yeast Saccharomyces paradoxus, we show that expression noise can be inferred in a reproducible manner.

11
BOSE: A Bayesian Order Statistics-Based Estimator for Recovering the Sample Mean and Standard Deviation

Pan, W.; Lu, Z.; Jiang, W.; Lim, J.; Xu, L.; Wang, X.

2026-07-01 bioinformatics 10.64898/2026.06.26.734829 medRxiv
Top 0.3%
1.1%
Show abstract

In meta-analyses of continuous outcomes, the sample mean and standard deviation (SD) are essential for synthesizing effect sizes across studies. However, clinical studies frequently report alternative summary statistics, such as the median, quartiles, and range. To enable inclusion of such studies, various methods have been proposed to estimate the sample mean and SD from these reported summaries. We propose the Bayesian Order Statistics-based Estimator (BOSE), which leverages the joint likelihood of observed order statistics together with weakly informative priors to obtain the full posterior distribution for the mean and SD without relying on computationally intensive iterative procedures such as Markov chain Monte Carlo algorithms. Our numerical studies demonstrate that BOSE performs competitively with existing approaches in estimating the mean, while achieving superior performance for estimating the SD across all evaluated scenarios, particularly in small-sample settings. Under non-normal distributions including skewed, heavy-tailed, and bimodal settings with mild or moderate deviations from normality, BOSE remains robust and stable, whereas methods specifically designed for skewed distributions may become unstable or even inapplicable. Beyond point estimation, BOSE naturally provides empirically validated posterior credible intervals, enabling researchers to formally quantify uncertainty for study-level estimates and make reliable, evidence-based decisions in meta-analytic research synthesis. A publicly accessible web application implementing BOSE and competing methods is also provided to facilitate practical use in meta-analytic research.

12
Simulation of synthetic health records for assessment of causal inference methods for vaccine efficacy

Velasco Pardo, V.; Daines, L.; Katikireddi, S. V.; Ritchie, L.; Robertson, C.; Simpson, C. R.; McCowan, C.; Swallow, B.

2026-07-19 infectious diseases 10.64898/2026.07.17.26358308 medRxiv
Top 0.3%
1.0%
Show abstract

Background During the COVID-19 pandemic, public health agencies used near real-time observational data to answer questions regarding vaccine effectiveness. However, traditional observational methods do not allow conclusions regarding counterfactual scenarios to be drawn from clinical data. Counterfactuals, which are outcomes that would have occurred under alternative interventions, can be used to formally assess the causal effects of public health interventions on health outcomes while accounting for the effects of confounding. Ideally individual patient data is used for the development of counterfactuals. Low-fidelity synthetic data may be useful for advancing methodological development where governance and privacy constraints prohibit access to sensitive personal data. Methods We simulated synthetic datasets based on the EAVE-II COVID-19 platform which has been limited to use for surveillance purposes. EAVE-II includes almost all resident people in Scotland registered with qualified general medical practitioners. Patient characteristics were simulated to reflect the known distribution of the Scottish population, accounting for dependencies between variables. Each synthetic dataset was encoded to different realistic scenarios for EAVEII 'ground truth' vaccine rollout and effectiveness results, explicitly stating the causal and confounding mechanisms, using a statistically sound method based on marginal structural models. Synthetic datasets of 100,000 individuals were then generated across five confounding scenarios and five severe outcome types. Results In scenarios with weak confounding, both unweighted and inverse probability of treatment weighted (IPTW) logistic regression recovered the true causal parameters. As confounding strength increased, only weighted models recovered the true mechanism. Conclusions Low-fidelity synthetic datasets simulated from EAVE-II data analysts to build and test causal inference pipelines, develop novel analysis pipelines, and train new researchers while awaiting access to real data. We showed how to generate synthetic datasets from a marginal structural model under different confounding scenarios.

13
Assessing tensor decomposition quality of immune profiling data from a dictionary learning perspective

Konstorum, A.; Xing, J.; Aeron, S.; Kilmer, M.; Kleinstein, S.

2026-07-09 bioinformatics 10.64898/2026.07.03.736447 medRxiv
Top 0.3%
0.9%
Show abstract

Systems-level immune profiling data arising from longitudinal studies of vaccination or infection has an inherent multi-index array structure. While tensor decomposition of such datasets has gained popularity, choosing a rank and trial for a decomposition is not straightforward. We show that taking into account the experimental data model can inspire the development of new metrics to assess the quality of a Non-negative CANDECOMP/PARAFAC (NCPD) decomposition, and can thus be used to choose a rank and trial for the decomposition. Moreover, we show how framing the results via a dictionary learning framework can better enable interpretation of the components of the decomposition.

14
GCBM-DCT-HV-Bio-NL-Grow-CHG-CSM-RHEC: A Unified Geometric, Biological, Causal, and Regenerative Framework for Mechanism-Aware Tissue and Connectome Modeling

Xu, T.; Hu, Z.; Sun, X.; Jin, L.; Xiong, M.

2026-06-29 bioinformatics 10.64898/2026.06.24.734320 medRxiv
Top 0.3%
0.9%
Show abstract

Modern biological prediction problems increasingly require models that go beyond Euclidean feature regression and local graph smoothing. Tissue, cellular, and connectome systems are nonlinear, geometry-dependent, intervention-sensitive, history-dependent, and subject to regenerative or homeostatic constraints. We propose GCBM/DCT/HV/Bio/NL/Grow/CHG/CSM/RHEC, a unified model for mechanism-aware biological prediction. The model integrates geometric connectome dynamics, differentiable charted tissue geometry, Hamiltonian latent transport, nonlinear biological kinetics, nested latent memory, continual growth without overwriting, causal hypergraph structure, causal structure modeling, and regenerative homeostatic error correction. Unlike Euclidean baselines, which treat observations as flat vectors, and local graph baselines, which use neighborhood smoothing without mechanistic structure, the proposed model represents biological states (Trapnell 2015) as coupled geometric, dynamical, causal, and regenerative objects. We evaluate the model on four synthetic toy studies, Toy A, B,C, D, designed to reflect increasing biological complexity: local Euclidean structure, nonlinear mechano-chemical interaction, causal intervention response, and out-of-distribution regenerative shift. Compared with Euclidean and local graph baselines, the full model achieves the lowest mean squared error across all four toy studies. Relative to the Euclidean baseline, the full model reduces MSE by approximately 63.0%, 89.1%, 89.0%, and 90.9% on Toy A, Toy B, Toy C, and Toy D, respectively. These results support the value of integrating geometry, mechanism, causal structure, adaptive growth, and regenerative correction into a single predictive architecture (Figure 1).

15
Genetic sensitivity analysis: estimating genetic confounding and environmentally mediated genetic effects using multiple exposures

Frach, L.; Rijsdijk, F.; Hannigan, L. J.; Dudbridge, F.; Pingault, J.-B.

2026-07-17 epidemiology 10.64898/2026.07.16.26358236 medRxiv
Top 0.3%
0.9%
Show abstract

Polygenic scores are imperfect measures of the additive genetic effects of common genetic variants. The resulting measurement error biases estimates of quantities of interest in epidemiological analyses integrating polygenic scores. For example, how much of an exposure-outcome association is genetically confounded can be substantially underestimated when using polygenic scores alone. Here we present extensions to Gsens, a genetic sensitivity analysis, which aims to correct for such measurement error using both polygenic scores and heritability estimates. Gsens now allows for multiple exposures and estimates several quantities of interest, i.e. genetic confounding, adjusted residual association (net of genetic confounding), genetic overlap and environmentally mediated genetic effects. We present derivations and simulations showing how Gsens accounts for measurement error in the polygenic score; we also show how estimation may be affected by misspecifications of the causal structure between exposures. Applying Gsens in the Norwegian Mother, Father and Child Cohort Study (MoBa), we uncover, among other results, substantial genetic confounding in the associations between multiple known risk factors for attention deficit hyperactivity disorder (ADHD), such as low birth weight and temperament, and measures of ADHD in childhood. The updated Gsens R package offers multiple options, including for missing data handling and customisable syntax. Our extended version of Gsens is applicable to a broad range of substantive questions in multiple disciplines.

16
GR-SAFS: A Graph-Regularized Stacking Framework with Adaptive Feature Selection for High-Dimensional Prognostic Biomarker Discovery

He, J.; Guan, J.

2026-06-28 bioinformatics 10.64898/2026.06.23.733986 medRxiv
Top 0.3%
0.8%
Show abstract

Identifying prognostic biomarkers from high-dimensional transcriptomic data poses a triple challenge: achieving sparsity, preserving biological network topology, and integrating complementary nonlinear signals. Existing methods typically ignore network structure, miss nonlinear interactions, or lack a principled mechanism to fuse heterogeneous model outputs. We introduce GR-SAFS (Graph-Regularized Stacking with Adaptive Feature Selection), a framework with three modules: a Graph-Lasso engine embedding gene co-expression network Laplacian priors, run in parallel with a Random Forest engine; an empirical cumulative distribution function (eCDF) alignment layer that places sparse and dense importances on a common percentile scale; and a diversity-penalized quadratic programming router whose strict convexity yields a unique global optimum. On the TCGA-LUAD cohort, GR-SAFS identifies a 20-gene signature with a training concordance index of 0.700. Across two independent crossplatform microarray cohorts, GR-SAFS is the only method whose frozen signature retains statistically significant risk stratification in every cohort, where stronger-C-index baselines lose significance on at least one external cohort. Functional enrichment anchors the signature to a coherent Wnt/{beta};-catenin axis. An open-source implementation is released for full reproducibility.

17
Two-Sample Instrumental Variables under Population Mismatch: A Transportability Framework with Bias Diagnostics

Qian, Y.; Song, Y.

2026-06-26 epidemiology 10.64898/2026.06.15.26355602 medRxiv
Top 0.4%
0.8%
Show abstract

Instrumental variable (IV) methods are widely used in health and social sciences to estimate causal treatment effects among compliers. In certain research settings, the instrument-treatment association (first stage) and the instrument-outcome association (reduced form) are each estimated from a different dataset. Two-Sample Instrumental Variables (TSIV), proposed by Angrist and Krueger (1992), addresses this by combining first-stage and reduced-form estimates from separate data sources into a single causal effect estimate. However, TSIV identification requires that instrument compliance behavior be consistent across the two samples, a condition that is rarely verified in practice. We show mathematically and empirically that when compliance differs between samples, the raw TSIV estimator does not converge to the true Local Average Treatment Effect (LATE) and instead attenuates toward a predictably biased limit proportional to the ratio of first-stage compliance rates between the two samples. To address this, we formalize a framework for estimating LATE with TSIV under two key assumptions: (1) Covariate Overlap, requiring that the two samples share sufficient common support in their covariate distributions, and (2) Compliance Transportability, requiring that compliance behavior is identical across populations after conditioning on observed covariates. We consider a setting in which a health policy instrument and outcomes are recorded in administrative claims while treatment and covariates are collected in a survey. We use a C-statistic derived from pooled covariates to detect population mismatch and an Inverse Probability Weighting (IPW) correction that reweights the first-stage sample to approximate the administrative covariate distribution. In Monte Carlo simulations across eight scenarios calibrated to a survey-Medicaid setting, IPW-TSIV reduces bias in estimating the LATE, achieving 88% reduction in the primary scenario, 82% under severe selection, and 79% when state-level expansion policy drives compliance heterogeneity. We further validate this framework using the Oregon Health Insurance Experiment, where partitioning the public-use lottery data (N = 24,646) into two non-overlapping samples with substantively meaningful compliance heterogeneity yields a verifiable benchmark against the true causal effect. IPW-TSIV reduces mean absolute bias by 71.6% relative to the oracle S2-specific LATE across 10 independent replications (C-statistic = 0.78), outperforms naive TSIV in all 10 splits, and reduces mean bias relative to the full-data LATE from +0.016 to +0.008. This framework provides applied researchers with actionable diagnostic thresholds to detect sample mismatch, validate transportability assumptions, and determine when structural TSIV estimation is reliable.

18
Neural Processes with Normalizing Flows for Wheat Height Estimation

Boss, M.;Volpi, M.;Roth, L.

2026-07-09 Plant Biology 10.64898/2026.06.24.734247 medRxiv
Top 0.4%
0.7%
Show abstract

In this work, we investigate modeling plant traits over time using neural processes, a class of machine learning models that learn distributions over functions. Plant growth is an inherently stochastic process with complex dynamics measured mostly at irregular times throughout the growing seasons. While individual trait trajectories may be simple, their distributions are shaped by complex interactions between genotype, environment, and other factors. In particular, we focus on plant height in wheat, a deceptively simple-looking trait with complex dynamics. To model these trajectory distributions, we evaluate neural processes and in particular extensions using normalizing flows, with different combinations of genotype and environmental covariates. For controlled evaluations, we generate synthetic wheat height trajectories calibrated against Swiss weather station records and the FIP1 dataset. To fully evaluate these trajectory distributions, we use signatures, vector representations of sequential data, together with Sig-MMD and the recently introduced CSig-MMD. Sig-MMD enables direct pathwise comparison of predicted and simulator trajectory distributions, while CSig-MMD focuses this comparison on the tail, including lodged trajectories. Together, these metrics allow us to assess whether the models capture the full distribution of growth trajectories, including rare outcomes.

19
Methodological guidelines for circadian modeling of Daylight Saving Time: application to the United States

Martin-Olalla, J. M.; Mira, J.

2026-06-22 public and global health 10.64898/2026.06.17.26355889 medRxiv
Top 0.4%
0.6%
Show abstract

Modeling the circadian impact of seasonal clock changing requires precise synchronization between solar and social time. This report critiques a recent study that associated disease prevalence in the United States with seasonal clock exposure. We identify a fundamental computational error in which a sign reversal of the longitudinal offset effectively inverted the US East-West axis, cross-correlating local health data with the circadian burden of hypothetical locations on the opposite side of a time zone. We outline the methodology for a correct modelization of the circadian process in the context of US geography.

20
Statistical Inference and Power Analysis for Comparative F1 and Fβ Scores under Correlated Classifier Pairs

Hsu, C.-Y.; Liu, Q.; Shyr, Y.

2026-07-17 dermatology 10.64898/2026.07.15.26358166 medRxiv
Top 0.4%
0.6%
Show abstract

As machine learning and artificial intelligence systems are increasingly used in healthcare, rigorous evaluation of their classification performance has become critical. The F1 and F{beta} scores are widely adopted metrics for assessing performance in imbalanced biomedical data. Recently, we introduced psF1, a unified statistical framework for inference and study design for single and comparative F1 and F{beta} scores under the assumption of independent classifiers. In practice, however, benchmarking two classifiers on the same dataset creates a correlated paired setting. Ignoring this intrinsic dependency leads to overestimation of the standard error and a substantial loss of statistical power. To address this, we develop psF1pair, an advanced framework for statistical inference and power analysis that explicitly accounts for correlations between classifier pairs. Extensive simulation studies demonstrate the performance of psF1pair, and its utility is further illustrated through application to a real-world imaging classification system. As expected, higher correlation between classifiers yields narrower confidence intervals and enhanced statistical power. A freely available R package is provided to facilitate implementation, supporting accurate evaluation and study design for predictive and classification models in biomedical research.