Back

Biostatistics

Oxford University Press (OUP)

Preprints posted in the last 30 days, ranked by how well they match Biostatistics's content profile, based on 24 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
An M-learner approach for heterogeneous mediation analysis with high-dimensional omics mediators

Li, X.; Wei, P.

2026-09-01 bioinformatics 10.64898/2026.08.25.747106 medRxiv
Top 0.1%
7.6%
Show abstract

Causal mediation analysis is widely used to identify biological pathways linking exposures to outcomes, but most methods assume homogeneous mediation effects across individuals. In high-dimensional omics settings, this assumption can mask important heterogeneity driven by demographic, genetic, or environmental factors. We propose the M-high-learner, a flexible framework for detecting heterogeneous mediation effects with high-dimensional mediators. The method identifies mediators with subgroup-specific indirect effects while distinguishing them from null or homogeneous signals and controlling the type I error rate. It is computationally efficient, scalable, and yields interpretable sub-types. Simulation studies show that the proposed approach achieves high power while maintaining accurate error control. Applications to the Framingham Heart Study and the Multi-Ethnic Study of Atherosclerosis reveal that the mediation role of gene expression in sexs effect on high-density lipoprotein varies across subgroups defined by body mass index and age. Our framework provides a practical tool for uncovering heterogeneous biological mechanisms in high-dimensional genomic studies. Author SummaryBiological processes linking risk factors to disease often differ across individuals, but many existing methods assume these processes are the same for everyone. This can hide important differences between groups. We developed a powerful method to identify when these pathways vary across subgroups using large-scale molecular data. Our approach detects differences in how intermediate biological factors contribute to outcomes in populations defined by characteristics such as age and body mass index. Applying our method to population studies, we found that some biological pathways operate differently across groups, suggesting that key mechanisms may be missed when differences are ignored. Our work provides a tool to better understand how disease-related processes vary across individuals, which may support more targeted and personalized approaches to health research.

2
Competing event regression on the relative subdistribution and cumulative-incidence scales

Mell, L. K.

2026-08-14 epidemiology 10.64898/2026.08.13.26360204 medRxiv
Top 0.1%
4.0%
Show abstract

In competing risks settings, covariate effects and group comparisons are usually assessed one event at a time - through log-rank or Cox tests on the cause-specific hazards, or Gray's test or Fine-Gray regression on a cumulative incidence function (CIF). This can obscure a clinically important quantity: the ratio between the event of interest and the competing event, since groups may differ little on the individual events yet differ sharply in their ratio. The generalized competing event (GCE) framework makes this ratio the object of inference; on the cause-specific scale the hazard ratio omega+(t) = lambda_1(t)/lambda_2(t) is estimated efficiently from a single stacked (Lunn-McNeil) model. We extend the framework to two scales that describe realized incidence. The subdistribution hazard ratio omega-tilde+(t) = lambda-tilde_1(t)/lambda-tilde_2(t) is estimated by a stacked, risk-set-weighted extension of the Lunn-McNeil construction; the cumulative-incidence ratio rho(t) = F_1(t)/F_2(t) - the odds that a subject's realized event by time t is the event of interest - by jackknife pseudo-observation regression of the Aalen-Johansen estimator. We relate the three contrasts: rho equals omega+ exactly under proportional cause-specific hazards, and equals omega-tilde+ only in the small-time limit under proportional subdistribution hazards, drifting toward 1 thereafter. The orthogonality that makes omega+ efficient is lost on both cumulative-incidence scales - omega tilde+ through overlapping weighted risk sets and shared censoring weights, rho through the shared all-cause survivor - so each carries a covariance term that must be handled and that bounds efficiency relative to the hazard-scale test. We derive the corresponding variances, study operating characteristics by simulation, illustrate on hypothetical prostate and head-and-neck cohorts, and provide an implementation in the gcemod R package.

3
Covariance Nonstationarity is Evident in Spatial Transcriptomics and Provides a New Categorization of Spatially Varying Genes

Velidi, P.; Wei, Z.; Nathoo, F.

2026-08-18 bioinformatics 10.64898/2026.08.10.743911 medRxiv
Top 0.1%
2.6%
Show abstract

BackgroundGaussian process models underlie many spatial transcriptomics tools but typically assume stationary covariance. While typically ignored, non-stationarity of spatial covariance in gene expression may correspond to tissue heterogeneity or cell aggregates. ResultsAcross 13 Visium datasets, we use approximate Bayes factors from R-INLA to compare stationary and non-stationary Matern covariance functions. Evidence for covariance non-stationarity appears in 3% to 50% of genes across tissue samples. We further characterize the power and false discovery rate of the Bayesian analysis of non-stationarity. We find that gene sets associated with immune, cytokine, and other effector functions are enriched among genes favoring non-stationary spatial covariance. ConclusionsCovariance stationarity is not a benign technical simplification in spatial transcriptomics; it is frequently violated, the violation is biologically structured, and it changes the definition and classification of spatially varying genes.

4
Mitigating the Effects of Population Stratification in Gene-Gene Interaction Studies

Das, N.; Ueki, M.

2026-08-21 genomics 10.64898/2026.08.18.745398 medRxiv
Top 0.1%
2.5%
Show abstract

Population stratification is a major source of inflated false positive rates in genome wide association studies. However, relatively few studies have examined its impact on gene-gene interaction detection, despite the importance of epistasis for understanding the genetic architecture of complex traits. In this study, we identify scenarios under which population stratification can inflate the interaction test statistics. Through analytical derivations and simulation studies, we show that this inflation is not adequately controlled by including principal components as covariates in the regression model. We then propose an alternative approach that effectively controls the inflation of false-positive rates for interaction test statistics due to population stratification by using single nucleotide polymorphism-by-population structure interaction as an additional covariate term in the regression model.

5
Design-informed Size Factor Estimation

Pocuca, T.; Pare, G.; Bolker, B. M.

2026-08-22 bioinformatics 10.64898/2026.08.13.744630 medRxiv
Top 0.1%
2.4%
Show abstract

Accurate normalization is essential for differential expression analysis of RNA-sequencing data. Popular normalization methods such as the median-of-ratios and trimmed mean of M-values do not leverage information from the experimental design. This may be inefficient in experiments with large-scale systematic expression changes or complex designs. Here, we introduce design-informed size factor estimation (disize), a normalization method that uses information from the experimental design to improve accuracy. disize uses a modified generalized linear mixed model to robustly distinguish between biological signal and sample-specific size factors. We also propose a mechanistically justified data-generating process for RNA-sequencing counts that is derived from previous models of transcription and sequencing. Through simulations based on this data-generating process and validating on true RNA-seq data, we show that disize recovers size factors more accurately than existing methods, particularly in challenging scenarios with low gene expression and a high proportion of differentially expressed genes; this in turn improves downstream analysis. disize provides a robust and accurate approach to normalization, highlighting the significant benefits of integrating experimental design information directly into normalization for transcriptomic datasets. Author summaryIn transcriptomic analysis, normalization adjusts for technical biases arising from library preparation and sequencing. Methods implemented in widely used packages like DESeq2 and edgeR ignore information in the experimental design during normalization. Incorporating information from the experimental design into a normalization method has the potential to yield more accurate results. To do this, we developed a new method, design-informed size factor estimation (disize), that uses a statistical model to jointly account for the biological signal defined by the design and the sample-specific batch effect. By separating the biological variation into its components, disize can more robustly estimate the batch effect. To validate our approach, we constructed a flexible simulation framework relying on a mechanistically justified data-generating process for RNA-seq data. Our benchmarks on both simulated and true RNA-seq data show that disize recovers the true size factors more accurately than existing methods, particularly in challenging scenarios with low counts or a high proportion of differentially expressed genes. This improved normalization yields more reliable downstream results in differential expression analysis.

6
RedFuMOS: A novel approach for multi-omics and clinical data-driven patient stratification

De Luca, S.; Fava, C.; Rizzo, G.; Visconti, A.; Berchialla, P.

2026-08-31 health informatics 10.64898/2026.08.26.26361415 medRxiv
Top 0.2%
1.7%
Show abstract

Background. Patient stratification from multi-omics and clinical data is essential for uncovering disease heterogeneity and moving toward more personalized treatment strategies. However, integrating heterogeneous data layers while identifying robust patient strata remains challenging. Methods. We introduce Reduced Fusion of Multi-Omics Stratification (RedFuMOS), a novel three-step approach for patient stratification based on mixed-type multi-omics data. RedFuMOS extends Similarity Network Fusion to accommodate mixed-type data layers and layer-specific similarity measures for data integration, includes a dimensionality reduction step to mitigate the curse of dimensionality, and performs patient stratification using density-based hierarchical clustering with HDBSCAN. It also implemented an automated optimization procedure to identify the best set of hyperparameters, minimizing the need for manual tuning. Results. RedFuMOS outperformed six state-of-the-art tools for multi-omics patient stratification in a comprehensive simulated benchmarking study, which also confirmed that, although computationally expensive, the dimensionality reduction step is crucial for achieving good stratification performance. Additionally, RedFuMOS identified two clinically relevant patient strata in a small real-world cohort of patients with Philadelphia chromosome-positive chronic myeloid leukaemia. Conclusion. RedFuMOS provides a flexible framework for integrating heterogeneous multi-omics and clinical data. RedFuMOS is available as an R package at http://github.com/delucasara/RedFuMOS.

7
Improving the Performance of Models Trained on Small EHR-Derived Samples by Leveraging External Data with Continual Learning Methods

Hui, J.; Xia, M.; Wilson, J.; Hill, E. D.; Scheer, A.; Franz, L.; Engelhard, M. M.; Goldstein, B. A.

2026-08-10 health informatics 10.64898/2026.08.08.26360010 medRxiv
Top 0.2%
1.7%
Show abstract

The performance of an EHR-based deep learning model trained on a small sample can be improved if more data is collected. Instead of collecting more data, the model can be trained on additional data from an analogous external source. However, this risks the model learning patterns in the external data that do not generalize to the target sample. Furthermore, data use agreements often prohibit combining datasets with medical records of different sources. We consider utilizing pre-existing methods in continual learning, namely the elastic weight consolidation (EWC) loss function and variational continual learning (VCL), both of which are regularization-based methods that we use to borrow external data and incorporate parameters from a model on external data into local model training. To investigate the utility of this modeling framework, we consider two binary classification tasks: (1) predicting which children will be diagnosed with autism spectrum disorder (ASD) from medical claims up to 18 months, and (2) predicting which patients with end-stage renal disease (ESRD) will be re-hospitalized within 30 days. Target datasets were derived from Duke University's EHR warehouse, and external datasets were sourced from either NC Medicaid claims for the ASD prediction task, or the United States Renal Data System (USRDS) for the rehospitalization prediction task. For both of these tasks, borrowing models - using either the EWC loss function or VCL - performed similarly to that of a model trained only on the full external data, when the sample size of target data used to train the model was small. That is, while a model that does not borrow using our methods performed poorly in low data regimes, the borrowing model instead matched the performance of a model trained on external data even when sample size of target data was small. In addition, an analysis of model predictions showed that models with small samples are better calibrated and more functionally similar to a model trained only on external data when the sample size is small.

8
A permutation-free family-wise error rate for the moderated top-gene scan under gene correlation

Dwyer, W. J.

2026-08-21 bioinformatics 10.64898/2026.08.17.745282 medRxiv
Top 0.2%
1.5%
Show abstract

A differential-expression scan reports the genes with the largest moderated t-statistics, so controlling the family-wise error rate means controlling the null distribution of the maximum statistic over genes. Under gene correlation this is widely believed to require permutation, because correlation changes the effective multiplicity and corrupts the empirical-Bayes variance prior behind the moderated t-statistic. We decompose that liberality by an error-budget ablation and show that, within the simulated model class, it reduces principally to an inflation of the empirical-Bayes prior degrees of freedom: substituting the true prior returns the family-wise error to the independent-gene small-sample baseline, so dependence imposes no separate barrier once the prior is correct. Correlation deflates the cross-gene spread of the log sample variances; because the prior degrees of freedom decreases in that spread, the prior is over-estimated and the moderated maximum turns liberal. Dividing the observed spread by one minus the mean squared gene correlation, estimated by a tuning-free spectral U-statistic with an unbiased trace target, reverses the mechanism and holds the family-wise error near the baseline at retained power without permutation. An observable instability index flags when severe co-expression should defer to permutation.

9
Synthetic Longitudinal Tabular Data Generation via Copula

Cai, H.; Yu, W.; Lu, R.; Chattopadhyay, I.; Zhang, X.; Liu, J.

2026-08-07 bioinformatics 10.64898/2026.08.03.742474 medRxiv
Top 0.2%
1.1%
Show abstract

Synthetic data generation is increasingly used to enable data sharing and secondary analysis while protecting participant privacy, particularly for longitudinal tabular health data, where repeated measures per subject create within-subject dependence that most synthetic data methods are not designed to preserve. Existing generative methods, particularly generative adversarial network (GAN)-based approaches, can model complex distributions, but their estimated dependence structures are often difficult to interpret and their performance may be unstable or prone to overfitting in modestly sized datasets. Here we show that eCDF-copula, a statistically rooted approach using the empirical cumulative distribution function (eCDF) and copula modeling, preserves within- and between-visit dependence structure. To handle pervasive missing data, we propose a two-stage strategy combining multiple imputation with copula-based synthesis, enabling a variance decomposition that quantifies replication variability across methods. We benchmarked the proposed approach against four established methods on two longitudinal clinical datasets spanning markedly different sample sizes (n = 120 vs. n = 3, 612). eCDF-copula achieved resemblance and utility exceeding those of state-of-the-art synthetic data methods, while maintaining comparable privacy.

10
Genetic Architecture and Sample Size Impact Relative Performance of Nonlinear Machine Learning and Standard Polygenic Risk Scores

Zhu, J.; Baousi, A.; Morris, A. P.; Guo, H.

2026-09-03 genetic and genomic medicine 10.64898/2026.08.29.26361109 medRxiv
Top 0.3%
1.1%
Show abstract

Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.

11
LRSPAT: A low-rank framework for spatial omics statistics

Frost, H. R.

2026-08-31 bioinformatics 10.64898/2026.08.26.747291 medRxiv
Top 0.3%
1.1%
Show abstract

We describe LRSPAT (low-rank spatial toolkit), a fast and memory-efficient framework for approximating measures of spatial association for high-dimensional data. While LRSPAT can be applied to any multivariate spatial dataset, development was motivated by the computational challenge of identifying spatially variable genes in high-resolution spatial transcriptomics (ST) data generated by technologies such as 10x Visium HD, Xenium and Atera. LRSPAT leverages a truncated SVD of the expression data and a thresholded spatial weights matrix to perform reduced-rank reconstruction of spatial statistics in the quadratic form family, including global and local versions of Moran's I, Geary's C, and Getis-Ord G. A regularization approach is leveraged to account for the inflated null distribution of spatial statistics computed on latent variables. By performing key operations on the low-dimensional embeddings, LRSPAT is orders of magnitude faster than standard implementations with significantly lower memory requirements. Because the low-rank approach denoises and desparsifies ST data, LRSPAT is also more accurate than standard techniques at identifying genes with true spatial expression patterns. The dramatic improvements in execution time and memory consumption enable the genome-wide analysis of spatially variable genes (SVGs) and exploration of the full range of hyperparameters including spatial scale, distance metric, and embedding rank. This preprint outlines the background and mathematical details of the approach with limited preliminary results and a short conclusion.

12
A software package for simple and rigorous survival machine learning analysis in biomedical research

Pybus, A.; Qiu, J.; Morais Lyra, P. C.; Dang, K.; Narvaez-Bandera, I.; Jolaogun, T.; Goecks, J.

2026-08-10 health informatics 10.64898/2026.08.05.26359034 medRxiv
Top 0.3%
1.0%
Show abstract

Survival analysis is a fundamental technique in biomedical research for modeling time-to-event data. It enables the identification of prognostic factors in disease, compares survival outcomes across treatment groups, and performs targeted treatment selection. A variety of machine learning (ML) approaches to survival analysis have emerged to complement classical statistical methods, especially for high-dimensional datasets with complex, nonlinear interactions between features. However, using survival ML methods requires addressing challenges such as censoring-unaware evaluation, overfitting, selecting performance metrics, and data leakage. To address these and other difficulties in using survival ML models, we developed the mlsurv software package. mlsurv is an open-source Python package built around three major design principles: 1) methodological rigor, including evidence-based model selection, leakage-free pipelines, and multi-metric evaluation, 2) multi-scale evaluation and interpretation, including population and subpopulation evaluation, patient-level explanations, and feature analysis, and 3) automated trust and transparency, including limitation flagging and TRIPOD+AI-aligned reporting. mlsurv bundles ten models spanning linear, ensemble, kernel, and deep learning families within a unified software package. We demonstrate mlsurv on the Chowell immunotherapy cohort (n=1,479). The survival-trained models achieve a test concordance index of 0.73 for overall survival prediction. Further, risk scores strongly correlate with the response-trained LORIS clinical score (|{rho}| up to 0.84), reflecting the overlap between prognostic and predictive signal. mlsurv enables biomedical researchers to conduct rigorous, multi-model survival analysis and benchmarking using minimal code with default best practices rather than implementing custom scripts and methodological safeguards from scratch.

13
Likelihood-Based Inference and Model Selection for Stochastic Gene Expression in Probability-Generating-Function Space

Wang, Y.; Shu, Z.; McAuley, K. B.; Cao, Z.

2026-08-25 systems biology 10.64898/2026.08.24.746673 medRxiv
Top 0.3%
1.0%
Show abstract

Selecting stochastic gene-expression models from single-cell counts requires accurate parameter inference and efficient model selection. Likelihood methods in count space can be costly when full stationary count distributions are unavailable, whereas approximate methods may lose accuracy. Probability generating functions (PGFs) offer a compact analytical alternative, but existing PGF workflows are generally not likelihood based and therefore rely on computationally intensive cross-validation. We develop a likelihood-based PGF framework for both tasks. Correlated empirical PGF values are used to construct a Gaussian quasi-likelihood for parameter inference and PGF-based Bayesian information criterion (BIC) for model selection. We show that the empirical PGF is exactly unbiased and that the parameter estimator is consistent, converges at the inverse-square-root sample-size rate, and is first-order asymptotically unbiased. For large samples and a uniquely preferred model, PGF-BIC selects the same model as leave-one-out cross-validation in PGF space.

14
Mechanistic Multi-Task Logistic Regression as an Alternative to Parametric Hazard Models in Joint Time-to-Event Analysis

Bisaso, K. R.; Kadada, K. R.; Bisaso, K. S.; Ette, E. I.

2026-08-18 pharmacology and therapeutics 10.64898/2026.08.15.26360512 medRxiv
Top 0.3%
0.9%
Show abstract

Background: Parametric time-to-event models require specification of a baseline hazard function, which may influence prediction when the underlying hazard shape is uncertain. This study compared conventional joint longitudinal time-to-event models with mechanistic Multi-Task Logistic Regression, which directly models the survival distribution without selecting a continuous parametric hazard family. Methods: A simulated dataset of 100 individuals with longitudinal sum of longest diameters and event outcomes was analyzed using a shared mechanistic tumor shrinkage regrowth model. Event submodels comprised exponential, Gompertz, Weibull, log-normal, log-logistic, and circadian hazards, mechanistic Multi-Task Logistic Regression, and a hybrid neural-mechanistic extension. All models were estimated jointly using shared patient-specific random effects and longitudinal data. Models were evaluated using longitudinal goodness-of-fit, visual predictive checks, five-fold cross-validated inverse-probability-of-censoring-weighted dynamic area under the curve and Brier scores, integrated Brier score, calibration, and event-interval negative log score. Results: Longitudinal parameter estimates and diagnostics were comparable across models. All conventional hazard models produced identical dynamic area under the curve values within prediction windows, although probabilistic accuracy differed. The log-normal hazard achieved the lowest overall integrated Brier score (0.1928). Mechanistic Multi-Task Logistic Regression achieved the highest later landmark discrimination (area under the curve 0.867 versus 0.798 for all hazard models) and the lowest mean event-interval negative log score (2.362). The hybrid model improved intermediate-landmark discrimination but not overall probabilistic accuracy. Conclusions: Mechanistic Multi-Task Logistic Regression provided competitive joint time-to-event prediction while avoiding baseline hazard-family selection. It represents a practical complementary alternative to parametric hazard modeling, particularly when hazard shape is uncertain and dynamic discrimination is important.

15
When the Background Matters: Topic-Dependent reference lists in GWAS and Exome Analyses

Timoney, B.; Guasoni, P.; Zade, K.; Bach, S.; Tropea, D.

2026-08-21 bioinformatics 10.64898/2026.08.14.744838 medRxiv
Top 0.3%
0.8%
Show abstract

Gene Ontology (GO) Biological Process overrepresentation analysis is widely used to interpret gene lists from genetic studies, yet results depend critically on the background (universe/reference list) against which enrichment is tested. This paper examines how genome-exome background mismatch alters GO Biological Process significance and induces annotation-driven bias. First, Monte Carlo simulations across multiple input gene list sizes show that enrichment p-values shift systematically when lists sampled from an exome-like universe are tested against a genome background (and vice versa), producing both inflation and deflation of significance depending on GO term composition; these shifts increase with gene list size. Second, applied analyses of gene lists derived from Genome-Wide Association Studies (GWAS) and Whole Exome Studies (WES) across brain, immune, and metabolic domains demonstrate that background choice changes the set of significant GO IDs, yielding reference-specific terms consistent with both Type I errors (false positives) and Type II errors (false negatives). Because genome backgrounds are commonly used by default, the practical risk is greatest when WES-derived lists are analyzed with genome reference lists. To support reproducible best practice, we provide a simple command set for selecting and documenting study-appropriate backgrounds and for assessing sensitivity of GO Biological Process results to the chosen universe.

16
bulk2scDiff: A Pseudobulk-Conditioned Diffusion Model for Bulk-to-Single-Cell RNASeq Generation

Xiao, J.; Raue, A.

2026-08-20 bioinformatics 10.64898/2026.08.20.745960 medRxiv
Top 0.3%
0.8%
Show abstract

Bulk RNA sequencing remains the predominant profiling strategy for large clinical cohorts, but it aggregates transcriptional signals across cell populations, thereby masking the underlying cellular heterogeneity. Inferring this heterogeneity from existing bulk transcriptomic data could extend large cohort-based studies that have already been profiled, but constitutes an underdetermined inverse problem, as one bulk profile can be compatible with multiple underlying cellular populations. Existing computational deconvolution methods address this problem primarily by estimating cell-type proportions or cell-type-averaged expression profiles rather than resolving expression at the level of individual cells. Here, we present bulk2scDiff, a proof-of-concept conditional diffusion framework that reformulates bulk-to-single-cell inference as conditional generation of single-cell expression profiles from pseudobulk transcriptomic input. We evaluated bulk2scDiff on two cancer single-cell RNA sequencing datasets, breast cancer and acute myeloid leukemia, where pseudobulk profiles were derived from the single-cell data and used as conditioning inputs, with the matched single-cell populations providing ground truth for controlled evaluation. Across both cases, bulk2scDiff closely reconstructed populations from training samples and generated biologically coherent single-cell populations for held-out samples, generalizing most consistently to recurrent immune features. A pseudobulk-swap control further confirmed sample-specific conditioning, with each sample corresponding pseudobulk yielding the closest agreement with its observed population in nearly all cases. Overall, our work establishes the feasibility of conditional diffusion for generating single-cell populations from pseudobulk transcriptomic profiles, providing a foundation for future evaluation with clinical bulk RNA sequencing data.

17
Sharing Aggregated Patient Counts in Place of Line-Level EHR Data: Analytic Fidelity and the Limits of Count Suppression for Privacy

Chen, Y.; McMurry, A.; Gottlieb, D.; Jones, J. R.; Strober, B. J.; Mandl, K. D.

2026-08-21 health informatics 10.64898/2026.08.18.26359984 medRxiv
Top 0.3%
0.8%
Show abstract

Objective. Privacy regulation constrains sharing line-level electronic health records (EHR) across institutions. One alternative is to aggregate counts into a cube, a table of counts for every combination of categorical variables, with cells below a threshold suppressed. This study asked whether common analyses on the cube reproduce conclusions from line-level data, and whether suppression prevents recovery of the small cells it is meant to hide. Materials and Methods. A Bayesian count-inference pipeline was built that reconstructs suppressed counts and doubles as a reconstruction attack. Applied to 285 pediatric kidney-transplant patients at Boston Children's Hospital, statistical fidelity (Jensen-Shannon divergence, Cramer's V, and R2) and analytical utility (marginal distributions, subgroup graft rejection odds ratios, and logistic-regression classification) were evaluated. Conditional Tabular GAN (CTGAN) synthetic data served as a comparator. Results. Statistical analyses on the cube recapitulated results from line-level data. Across 106 demographic-by-medication subgroups, a bootstrap mean of 3.5 subgroups showed a significant graft-rejection association. The cube's odds-ratio sign changes reversed no significant associations, versus 2.3 for CTGAN. The same reconstruction also defeated suppression: in a 10-variable cube, 76.6% of suppressed cube cells were recovered exactly (14,554 of 18,994), including 85.5% of single-patient cells. Discussion. The cube reproduced common kidney-transplant analyses, but the same reconstruction also recovered suppressed cells; fidelity and privacy risk are thus two faces of one reconstruction rather than independent properties. Conclusions. The cube is a useful surrogate for these kidney-transplant analyses only when paired with a stronger privacy mechanism. This study demonstrated reconstructability of suppressed counts, not re-identification.

18
MOFTy: Multimodal Gaussian Process Factor Analysis with Numerical Information Field Theory

Neumann, M.; Arras, P.; Kaster, A.-K.; Ott, A.

2026-08-21 bioinformatics 10.64898/2026.08.12.744240 medRxiv
Top 0.4%
0.6%
Show abstract

Multimodal Gaussian process factor analysis provides a flexible framework for dimensionality reduction in temporally or spatially resolved omics data. Existing approaches, however, typically rely on pre-specified Gaussian process kernel families and do not explicitly separate each latent factor into a component capturing gradual, smooth variation and a complementary component capturing fine-scale, non-smooth variation. Here, we present MOFTy, a Bayesian multimodal factor analysis framework based on numerical information field theory (NIFTy) that replaces fixed kernel families with the flexible correlated field model in NIFTy and enables explicit additive component separation within each latent factor with quantified uncertainty. NIFTy has been successfully applied to high-resolution Bayesian imaging in astrophysics and facilitates scalable, curvature-aware variational inference for efficient posterior approximations. We validate MOFTy on simulated data; applications to published multi-omics data demonstrate that MOFTy disentangles latent spatial structures by separating smooth gradients from localized fine-scale heterogeneity in human glioblastoma and recovers cross-modal patterns in a mouse gastrulation dataset.

19
A statistical framework for disease classification with scRNA-Seq Data

Xiao, Z.; Torous, W.; Cheng, J.; Cho, R.; Purdom, E.

2026-08-25 bioinformatics 10.64898/2026.08.21.746294 medRxiv
Top 0.4%
0.6%
Show abstract

Motivation Bulk RNA-sequencing based disease classification obscures cell-type specific signals by aggregating gene expression across heterogeneous tissues. Although single-cell RNA-seq tackles this limitation, summarizing and deriving patient-level predictors while retaining biological interpretability remains challenging. Standard sparse methods, such as lasso, often select arbitrary scattered gene sets without leveraging the underlying cell type structures revealed by single-cell data. Results: We introduce a two-stage statistical framework for interpretable patient-level disease classification from single-cell data. We first construct a gene-by-cell-type pseudobulk matrix that summarize single-cell expression for each patient. We then fit a multinomial logistic regression model with sparse group lasso penalty, inducing sparsity at both the cell type and gene levels. Across datasets of systemic lupus erythematosus, COVID-19, and colorectal cancer, our framework either matched or outperformed lasso and random forest baselines. Importantly, our models recovered biologically coherent, cell-type specific gene signatures consistent with known disease mechanisms, demonstrating improved interpretability without sacrificing predictive accuracy. Availability: The scSGL R package implementing the Sparse Group Lasso classification framework described in this paper is available at https://github.com/zhiweixiao/scSGL (version 0.99.1). Code to reproduce the actual cross-validation, model fitting, and prediction analyses on the three datasets reported here is available at https://github.com/zhiweixiao/scSGL-manuscript.

20
Constructing microbiome co-occurrence networks with confidence: A conditional, nonparametric, inference-based approach

Song, H.; Xiang, Y.; Liu, H.; Ling, W.; Plantinga, A. M.; Srinivasan, S.; Dun, Y.; Zhao, N.; Sun, S.; Engel, S. M.; Simon, N.; Wu, M. C.

2026-09-01 bioinformatics 10.64898/2026.08.27.747483 medRxiv
Top 0.4%
0.6%
Show abstract

Constructing microbial association networks is a common strategy for exploring relationships among taxa in microbiome studies. Although marginal correlation methods are easy to implement and allow formal inference, they can produce spurious edges driven by indirect associations through other taxa. Conditional graphical-modeling methods aim to recover direct associations, but many rely on Gaussian or linear assumptions and often provide limited uncertainty quantification. We propose a conditional, nonparametric approach based on the scaled expected conditional covariance (SEcov). SEcov measures population-level conditional association by residualizing each taxon with respect to the remaining taxa and scaling the resulting expected conditional covariance. The resulting estimator can incorporate flexible machine-learning methods for conditional-mean estimation and admits asymptotic normal inference, enabling p-values and confidence intervals for taxon-pair associations. We demonstrate through simulation studies that our proposed approach improves network recovery relative to other methods, and we illustrate the new method via construction of a co-occurrence network for the vaginal microbiome during pregnancy. IMPORTANCEHigh-throughput sequencing has made it possible to characterize microbial communities at large scale, and network analysis is widely used to summarize relationships among taxa. However, networks based on marginal correlations may include indirect associations, whereas many conditional graphical models rely on assumptions that may be difficult to justify for sparse, zero-inflated, compositional microbiome data. SEcov offers a practical alternative by estimating conditional associations nonparametrically and attaching inferential uncertainty to individual edges. This allows investigators to construct microbiome networks using statistically interpretable evidence for taxon-pair associations, rather than relying solely on arbitrary correlation cutoffs or regularization tuning parameters.