Back

Biometrics

Oxford University Press (OUP)

Preprints posted in the last 30 days, ranked by how well they match Biometrics's content profile, based on 23 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
Competing event regression on the relative subdistribution and cumulative-incidence scales

Mell, L. K.

2026-08-14 epidemiology 10.64898/2026.08.13.26360204 medRxiv
Top 0.1%
12.5%
Show abstract

In competing risks settings, covariate effects and group comparisons are usually assessed one event at a time - through log-rank or Cox tests on the cause-specific hazards, or Gray's test or Fine-Gray regression on a cumulative incidence function (CIF). This can obscure a clinically important quantity: the ratio between the event of interest and the competing event, since groups may differ little on the individual events yet differ sharply in their ratio. The generalized competing event (GCE) framework makes this ratio the object of inference; on the cause-specific scale the hazard ratio omega+(t) = lambda_1(t)/lambda_2(t) is estimated efficiently from a single stacked (Lunn-McNeil) model. We extend the framework to two scales that describe realized incidence. The subdistribution hazard ratio omega-tilde+(t) = lambda-tilde_1(t)/lambda-tilde_2(t) is estimated by a stacked, risk-set-weighted extension of the Lunn-McNeil construction; the cumulative-incidence ratio rho(t) = F_1(t)/F_2(t) - the odds that a subject's realized event by time t is the event of interest - by jackknife pseudo-observation regression of the Aalen-Johansen estimator. We relate the three contrasts: rho equals omega+ exactly under proportional cause-specific hazards, and equals omega-tilde+ only in the small-time limit under proportional subdistribution hazards, drifting toward 1 thereafter. The orthogonality that makes omega+ efficient is lost on both cumulative-incidence scales - omega tilde+ through overlapping weighted risk sets and shared censoring weights, rho through the shared all-cause survivor - so each carries a covariance term that must be handled and that bounds efficiency relative to the hazard-scale test. We derive the corresponding variances, study operating characteristics by simulation, illustrate on hypothetical prostate and head-and-neck cohorts, and provide an implementation in the gcemod R package.

2
Constructing microbiome co-occurrence networks with confidence: A conditional, nonparametric, inference-based approach

Song, H.; Xiang, Y.; Liu, H.; Ling, W.; Plantinga, A. M.; Srinivasan, S.; Dun, Y.; Zhao, N.; Sun, S.; Engel, S. M.; Simon, N.; Wu, M. C.

2026-09-01 bioinformatics 10.64898/2026.08.27.747483 medRxiv
Top 0.1%
6.8%
Show abstract

Constructing microbial association networks is a common strategy for exploring relationships among taxa in microbiome studies. Although marginal correlation methods are easy to implement and allow formal inference, they can produce spurious edges driven by indirect associations through other taxa. Conditional graphical-modeling methods aim to recover direct associations, but many rely on Gaussian or linear assumptions and often provide limited uncertainty quantification. We propose a conditional, nonparametric approach based on the scaled expected conditional covariance (SEcov). SEcov measures population-level conditional association by residualizing each taxon with respect to the remaining taxa and scaling the resulting expected conditional covariance. The resulting estimator can incorporate flexible machine-learning methods for conditional-mean estimation and admits asymptotic normal inference, enabling p-values and confidence intervals for taxon-pair associations. We demonstrate through simulation studies that our proposed approach improves network recovery relative to other methods, and we illustrate the new method via construction of a co-occurrence network for the vaginal microbiome during pregnancy. IMPORTANCEHigh-throughput sequencing has made it possible to characterize microbial communities at large scale, and network analysis is widely used to summarize relationships among taxa. However, networks based on marginal correlations may include indirect associations, whereas many conditional graphical models rely on assumptions that may be difficult to justify for sparse, zero-inflated, compositional microbiome data. SEcov offers a practical alternative by estimating conditional associations nonparametrically and attaching inferential uncertainty to individual edges. This allows investigators to construct microbiome networks using statistically interpretable evidence for taxon-pair associations, rather than relying solely on arbitrary correlation cutoffs or regularization tuning parameters.

3
Design-informed Size Factor Estimation

Pocuca, T.; Pare, G.; Bolker, B. M.

2026-08-22 bioinformatics 10.64898/2026.08.13.744630 medRxiv
Top 0.1%
4.3%
Show abstract

Accurate normalization is essential for differential expression analysis of RNA-sequencing data. Popular normalization methods such as the median-of-ratios and trimmed mean of M-values do not leverage information from the experimental design. This may be inefficient in experiments with large-scale systematic expression changes or complex designs. Here, we introduce design-informed size factor estimation (disize), a normalization method that uses information from the experimental design to improve accuracy. disize uses a modified generalized linear mixed model to robustly distinguish between biological signal and sample-specific size factors. We also propose a mechanistically justified data-generating process for RNA-sequencing counts that is derived from previous models of transcription and sequencing. Through simulations based on this data-generating process and validating on true RNA-seq data, we show that disize recovers size factors more accurately than existing methods, particularly in challenging scenarios with low gene expression and a high proportion of differentially expressed genes; this in turn improves downstream analysis. disize provides a robust and accurate approach to normalization, highlighting the significant benefits of integrating experimental design information directly into normalization for transcriptomic datasets. Author summaryIn transcriptomic analysis, normalization adjusts for technical biases arising from library preparation and sequencing. Methods implemented in widely used packages like DESeq2 and edgeR ignore information in the experimental design during normalization. Incorporating information from the experimental design into a normalization method has the potential to yield more accurate results. To do this, we developed a new method, design-informed size factor estimation (disize), that uses a statistical model to jointly account for the biological signal defined by the design and the sample-specific batch effect. By separating the biological variation into its components, disize can more robustly estimate the batch effect. To validate our approach, we constructed a flexible simulation framework relying on a mechanistically justified data-generating process for RNA-seq data. Our benchmarks on both simulated and true RNA-seq data show that disize recovers the true size factors more accurately than existing methods, particularly in challenging scenarios with low counts or a high proportion of differentially expressed genes. This improved normalization yields more reliable downstream results in differential expression analysis.

4
MeshScope-Scenario: A Seeded Monte Carlo Framework for Probabilistic Assessment of ICU and HCU Capacity Shortfall in Japan's Secondary Medical Areas, Incorporating Inter-Zone Transfer and Seasonal Surge

Ohno, K.; Hirai, M.; Hashimoto, S.

2026-08-10 health systems and quality improvement 10.64898/2026.08.05.26359832 medRxiv
Top 0.1%
3.3%
Show abstract

Background: Descriptive mapping of intensive care unit (ICU) and high care unit (HCU) capacity across Japan's secondary medical areas (SMAs) characterizes where beds exist, but medical planning also requires answers to prospective questions: how likely is a capacity shortfall under demand surge, which assumptions drive that risk, and how much protection do inter-zone transfer arrangements provide. No openly available tool addresses these questions at the SMA level, the geographic unit at which Japanese medical plans are written. Methods: We developed MeshScope-Scenario, a probabilistic capacity-demand framework operating on the MeshScope-Region platform. For a selected SMA, observed inputs (notified ICU/HCU beds from the Hospital Bed Function Reports; resident population) are combined with four explicitly flagged assumption parameters - effective staffed-bed rate, concurrent severe-care demand per 100,000 population, surge multiplier, and net cross-boundary inflow - each with a user-specified distribution. A seeded Monte Carlo engine (deterministic reproduction under a fixed seed) estimates the distribution of bed shortfall; interventions are compared under common random numbers. Parameter dependence is introduced by a Gaussian copula with automatic positive-semidefinite correction; global sensitivity is quantified by Sobol first-order and total-order indices (Saltelli sampling, Jansen estimators) alongside a deterministic one-at-a-time tornado analysis. A two-zone extension transfers unmet demand to the nearest ICU-holding SMA using road-network travel times measured in MeshScope-Region, with a transfer time limit and an acceptance cap; because both zones share the same systemic draws, correlated exhaustion of donor capacity under surge ("shared-fate" risk) is represented structurally. A seasonal layer applies twelve monthly surge multipliers and reports the distribution of annual maximum shortfall and month-specific shortfall probabilities. The engine is a dependency-free pure-function module verified by 40 statistical tests. Results: The framework reproduces identical output under identical seed and input; a flat seasonal profile reproduces the non-seasonal model exactly; copula factorization error is below 1e-9; and 10,000 iterations across three intervention variants complete in approximately 50 ms in a standard browser, permitting fully interactive use. Applied to three archetypal SMAs from the observed FY2024 supply map (seed 42, 20,000 iterations, demand prior 5 per 100,000), an ICU-zero zone with a transfer partner 48 minutes away has shortfall probability 66.0% (P50 4.8, P90 20.7 beds); a median metropolitan zone, 32.7% (P90 5.9), with Sobol indices ranking demand density and surge dominant; a high-supply zone shows zero shortfall up to approximately 2.5x surge. For the ICU-zero zone, independent-donor reasoning credits the transfer arrangement with a 2.27-bed reduction in expected shortfall, of which shared-fate correlation removes 93%; the probability of severe shortfall under the arrangement equals that with no arrangement at all, while a committed pool of five donor beds (6% of donor effective supply) restores a 13-point reduction and more than halves the arrangement's correlation exposure. Conclusions: MeshScope-Scenario extends SMA-level capacity mapping from description to prospective risk assessment. All demand-side inputs are declared assumptions with adjustable distributions rather than estimates presented as fact; the framework's value is to make the consequences of those assumptions, and their interaction with observed supply, explicit, reproducible, and inspectable for planning deliberation.

5
An M-learner approach for heterogeneous mediation analysis with high-dimensional omics mediators

Li, X.; Wei, P.

2026-09-01 bioinformatics 10.64898/2026.08.25.747106 medRxiv
Top 0.1%
3.1%
Show abstract

Causal mediation analysis is widely used to identify biological pathways linking exposures to outcomes, but most methods assume homogeneous mediation effects across individuals. In high-dimensional omics settings, this assumption can mask important heterogeneity driven by demographic, genetic, or environmental factors. We propose the M-high-learner, a flexible framework for detecting heterogeneous mediation effects with high-dimensional mediators. The method identifies mediators with subgroup-specific indirect effects while distinguishing them from null or homogeneous signals and controlling the type I error rate. It is computationally efficient, scalable, and yields interpretable sub-types. Simulation studies show that the proposed approach achieves high power while maintaining accurate error control. Applications to the Framingham Heart Study and the Multi-Ethnic Study of Atherosclerosis reveal that the mediation role of gene expression in sexs effect on high-density lipoprotein varies across subgroups defined by body mass index and age. Our framework provides a practical tool for uncovering heterogeneous biological mechanisms in high-dimensional genomic studies. Author SummaryBiological processes linking risk factors to disease often differ across individuals, but many existing methods assume these processes are the same for everyone. This can hide important differences between groups. We developed a powerful method to identify when these pathways vary across subgroups using large-scale molecular data. Our approach detects differences in how intermediate biological factors contribute to outcomes in populations defined by characteristics such as age and body mass index. Applying our method to population studies, we found that some biological pathways operate differently across groups, suggesting that key mechanisms may be missed when differences are ignored. Our work provides a tool to better understand how disease-related processes vary across individuals, which may support more targeted and personalized approaches to health research.

6
Covariance Nonstationarity is Evident in Spatial Transcriptomics and Provides a New Categorization of Spatially Varying Genes

Velidi, P.; Wei, Z.; Nathoo, F.

2026-08-18 bioinformatics 10.64898/2026.08.10.743911 medRxiv
Top 0.1%
2.4%
Show abstract

BackgroundGaussian process models underlie many spatial transcriptomics tools but typically assume stationary covariance. While typically ignored, non-stationarity of spatial covariance in gene expression may correspond to tissue heterogeneity or cell aggregates. ResultsAcross 13 Visium datasets, we use approximate Bayes factors from R-INLA to compare stationary and non-stationary Matern covariance functions. Evidence for covariance non-stationarity appears in 3% to 50% of genes across tissue samples. We further characterize the power and false discovery rate of the Bayesian analysis of non-stationarity. We find that gene sets associated with immune, cytokine, and other effector functions are enriched among genes favoring non-stationary spatial covariance. ConclusionsCovariance stationarity is not a benign technical simplification in spatial transcriptomics; it is frequently violated, the violation is biologically structured, and it changes the definition and classification of spatially varying genes.

7
STREAM-SMR: Sequential Bayesian state-space monitoring of standardized mortality ratios - a simulation comparison with risk-adjusted CUSUM and EWMA

Ohno, K.

2026-08-19 intensive care and critical care medicine 10.64898/2026.08.18.26360662 medRxiv
Top 0.1%
2.3%
Show abstract

Background. Sequential monitoring of risk-adjusted mortality in intensive care typically relies on alarm-generating control charts - the risk-adjusted CUSUM, EWMA, or VLAD. These charts signal deterioration but do not return what clinicians and registry stewards ultimately need to interpret: a calibrated, continuously updated estimate of the standardized mortality ratio (SMR) itself. Methods. STREAM-SMR is a conjugate gamma-Poisson dynamic generalized linear model in which the latent log-SMR evolves through a discount factor delta and the alarm statistic is the posterior exceedance probability P(SMR > 1). The construction is closed-form, exact for zero-death months, and computationally trivial at registry scale. Under a protocol frozen before any evaluation runs and calibrated to published national ICU registry aggregates, we compared STREAM-SMR (delta in {0.90, 0.95, 0.97}) with the risk-adjusted CUSUM and risk-adjusted EWMA across three facility-volume strata (50, 200, and 800 annual admissions) and five change scenarios (sustained steps, gradual drift, transient deterioration, and improvement). All methods were Monte-Carlo-calibrated to a common 5% false-alarm probability over a 60-month horizon, with 1,000 replications per cell. Results. At delta = 0.90, STREAM-SMR matched the detection frontier of the risk-adjusted CUSUM to within one to two months across sustained-shift scenarios - median delay for an SMR step to 1.5 of 6 versus 5 months in large facilities, 14 versus 14 in medium, and 20 versus 22 in small - while returning filtered SMR estimates whose 95% credible intervals held at least 91% empirical coverage in every scenario-stratum cell. Empirical false-alarm probabilities were close to the 5% nominal target for all methods (range 0.035-0.066). For a three-month transient deterioration, detection was faster with STREAM-SMR conditional on occurring, but overall detection probability favored the CUSUM in large facilities. Conclusions. STREAM-SMR unifies monitoring and estimation in a single Bayesian object: for a detection-delay premium of at most one to two months against the theoretically optimal CUSUM, it returns an interpretable, uncertainty-quantified SMR trajectory at every time point. The discount factor is an explicit dial between estimate smoothness and detection speed. Simulation code and the frozen protocol are publicly archived.

8
A Simulation Study Comparing Multiple Imputation and Complete Case Analysis for Handling Missing Preschool Body Mass Index

Savu, A.; Dover, D. C.; Hajihosseini, M.; Gaudet, L. A.; Kaul, P.

2026-08-14 epidemiology 10.64898/2026.08.13.26360115 medRxiv
Top 0.1%
2.1%
Show abstract

Background and Objective. Missing data frequently occurs in health databases and can bias analyses if not correctly dealt with. Using real-world data, we compared complete-case and multiple-imputation methods for recovering true parameters of a multivariable logistic regression model for the association between maternal glucose levels during pregnancy and child excess weight at preschool age, where missing values were present in as much as 30% of our sample. Methods. This study utilized a cohort of 130,424 children with complete preschool-age body mass index (BMI) measurements from the Calgary and Edmonton health regions of Alberta, Canada. In the complete BMI data, we introduced missingness through deletion following three distinct mechanisms: missing completely at random (MCAR), at random (MAR), and not at random (MNAR). To handle the missing data created, we employed complete-case and multiple-imputation methods. Maternal glucose levels during pregnancy were categorized into five groups and its association with child excess weight at pre-school age was determined based on a logistic regression model using the full observed data (yielding true values), observed data that was not deleted (complete-case estimates), and imputed data (multiple-imputation estimates). The accuracy of complete-case and multiple-imputation estimates were evaluated against the true values. Finally, we conducted a sensitivity analysis for the MNAR mechanism using pattern-mixture models with an additive shift. Results. Under MCAR and MAR, multiple-imputation generally outperformed complete-case, yielding smaller absolute and relative bias. Both methods achieved high significance ([≥] 0.96) for most effects. Mean squared errors for multiple-imputation and complete-case were similar missing completely at random, missing at random, and coverage was consistently high ([≥] 0.99). Under MNAR, both complete-case and multiple-imputation showed poor performance regarding bias and statistical significance. Sensitivity analysis using pattern-mixture models indicated performance varied by specific effect. Conclusions. Under MCAR and MAR, multiple-imputation introduced higher bias but demonstrated superior overall performance based on mean squared error and restored statistical power. Conversely, both methods failed under MNAR, where pattern-mixture modeling sensitivity analyses revealed highly variable, effect-specific performance due to unverifiable shift assumptions. When faced with missing data, researchers should assess missingness mechanisms, report both complete-case and multiple-imputation estimates under MCAR/MAR while accounting for power-versus-bias tradeoffs, and employ pattern-mixture sensitivity analyses to test robustness when MNAR is plausible.

9
Mitigating the Effects of Population Stratification in Gene-Gene Interaction Studies

Das, N.; Ueki, M.

2026-08-21 genomics 10.64898/2026.08.18.745398 medRxiv
Top 0.1%
1.9%
Show abstract

Population stratification is a major source of inflated false positive rates in genome wide association studies. However, relatively few studies have examined its impact on gene-gene interaction detection, despite the importance of epistasis for understanding the genetic architecture of complex traits. In this study, we identify scenarios under which population stratification can inflate the interaction test statistics. Through analytical derivations and simulation studies, we show that this inflation is not adequately controlled by including principal components as covariates in the regression model. We then propose an alternative approach that effectively controls the inflation of false-positive rates for interaction test statistics due to population stratification by using single nucleotide polymorphism-by-population structure interaction as an additional covariate term in the regression model.

10
Describing health inequalities without distortion: Simple-Means MAIHDA vs Random-Effects MAIHDA

Merlo, J.; Bashir, N. Z.; Rodriguez-Lopez, M.; Khalaf, K.; Öberg, J.; Perez-Vicente, R.

2026-08-18 epidemiology 10.64898/2026.08.17.26360592 medRxiv
Top 0.2%
1.5%
Show abstract

Multilevel Analysis of Individual Heterogeneity and Discriminatory Accuracy (MAIHDA) describes health inequalities through three components: (i) specific contextual effects (SCE), (ii) general contextual effects (GCE), and (iii) discriminatory accuracy of the context. We present Simple-Means MAIHDA (S-MAIHDA), which estimates each stratum directly from its observed individuals, with no distributional assumption. The observed proportions are unbiased whatever the stratum size, and their confidence intervals report the uncertainty honestly. S-MAIHDA operationalises the three components on the probability scale. The SCE are the raw and standardised stratum prevalences and the modification of the sociodemographic average differences by the area. The GCE are the variance partition coefficient (VPC) and the contextual structuring of the between-stratum inequality, expressed as the contextual clustering of inequalities, the additive sociodemographic differences, and the contextual modification of inequalities (CMI). The contextual discriminatory accuracy is expressed by the area under the ROC curve (AUC), and the sensitivity and specificity at the population prevalence as the threshold for a possible intervention. Because its estimates are the observed data themselves, S-MAIHDA is the canonical description, and the compare diagnostic quantifies how Random-Effects MAIHDA (RE-MAIHDA), the usual implementation, departs from it: RE shrinkage pulls small strata towards the overall mean and can hide the very inequalities the analysis seeks. The approach is implemented in the smaihda Stata command and reproduced in free Python code. We illustrate S-MAIHDA on register data from Malmo, Sweden (43,291 individuals; 300 area-sociodemographic strata), showing how the three components separate two contrasting outcomes: psychotropic medication use, almost purely sociodemographic, stable across areas, with weak contextual structuring (VPC {approx} 4%, CMI {approx} 0%); and choice of a private general practitioner, strongly geographical (VPC {approx} 11%, CMI {approx} 17%), with the sociodemographic differences reshaped and amplified in wealthy areas. RE-MAIHDA attenuated inequalities. For describing inequalities, S-MAIHDA preserves what the data show.

11
MASCOT-DS improves transmission dynamics inference by integrating multiple epidemiological data streams with phylodynamic inference

Weidemueller, P. H.; Esquivel Gomez, L. R.; Rodriguez-Barraquer, I.; Mueller, N. F.

2026-08-25 epidemiology 10.64898/2026.08.21.26361056 medRxiv
Top 0.2%
1.5%
Show abstract

Tracking how an infectious disease spreads in time and space relies on several distinct sources of surveillance data, reported case counts, viral concentrations in wastewater, seroprevalence surveys, and pathogen genomic sequences, each of which is imperfect and captures only part of the underlying transmission process. These data streams are typically analyzed separately or with highly parameterized, disease-specific models, making it difficult to combine their complementary strengths. Here we present MASCOT-DataStreams (MASCOT-DS), a BEAST2 software package that extends the structured coalescent model MASCOT to jointly infer prevalence over time and transmission rates between locations from any combination of case counts, wastewater concentrations, seroprevalence surveys, and pathogen phylogenies. Using simulated outbreaks in structured populations, we show that MASCOT-DS accurately recovers true prevalence trajectories and between-location migration rates. We then apply MASCOT-DS to genomic, case count, wastewater, and seroprevalence data from the SARS-CoV-2 Epsilon wave (winter 2020-21) in three San Francisco Bay Area counties, reconstructing county-level prevalence dynamics and quantifying transmission within and into the region. By systematically removing individual data streams, we find that genomic data are uniquely required to estimate transmission between locations, while seroprevalence data are essential for anchoring the overall magnitude of an outbreak; case counts and wastewater concentrations play largely interchangeable roles in capturing outbreak shape. These results demonstrate that integrating complementary epidemiological data streams substantially increases the certainty of transmission dynamics estimates compared to relying on any single data stream, and provides a framework for evaluating the added value of different surveillance strategies.

12
Relevance Based Prediction: A Transparent, Non-Artificial Intelligence, Mathematical Solution to Personalized Opioid Treatment

Robinson, C. L.; Turkington, D.; Lee, L.; Kritzman, M.; Yong, R. J.

2026-08-10 pain medicine 10.64898/2026.08.07.26359966 medRxiv
Top 0.2%
1.4%
Show abstract

Accurate prediction of individual medical outcomes is essential for optimizing treatment allocation amid rising costs, coverage denials, and limited clinical resources. Traditional predictive models, including regression and neural networks, rely on average effects and cannot tailor predictions to the specific circumstances of individual cases. We present relevance-based prediction (RBP), a model-free method that predicts outcomes as weighted averages of observed cases, with weights determined by a rigorously defined measure of relevance. Unlike model-based methods that rely on fixed calibrated parameters, RBP revisits the original data for each prediction and customizes both the cases and variables used. Applied to opioid treatment, RBP provides case-specific insights unavailable from conventional models, including how each prior case informs a prediction, how each variable affects its reliability and value, and how reliable the prediction is before it is made. These individualized insights may prevent misleading average-based decisions and reduce harmful or suboptimal treatment.

13
Increasing Lung Cancer Screening Participation Using an Informational Video Nudge: A Randomized Feasibility Trial

Wain, K. F.; Carroll, N. M.; Maclennan, A. J.; Hixon, B.; Steiner, J.; Ritzwoller, D. P.

2026-09-01 health systems and quality improvement 10.64898/2026.08.28.26361654 medRxiv
Top 0.2%
1.4%
Show abstract

Purpose: Lung cancer screening (LCS) with low-dose computed tomography (LDCT) reduces lung cancer mortality, yet screening participation remains low. We evaluated whether a brief informational video nudge delivered immediately before a scheduled clinical encounter increased LCS ordering and baseline LCS completion. Patients and Methods: We conducted a randomized feasibility trial within Kaiser Permanente Colorado from March through October 2025. LCS-eligible patients with an upcoming primary care or pulmonology appointment were assigned to intervention or usual care based on birth month. Intervention patients were split into two group, a group who received the LCS informational video nudge via text message within 24 hours of an eligible appointment; and second group who received the text plus a QR code video link during appointment rooming. Outcomes included LCS orders, baseline LCS-LDCT completion, and video engagement. Multivariable logistic regression was used to evaluate factors associated with LCS ordering. Results: Among 1,093 patients, 549 were assigned to intervention and 544 to usual care. Intervention patients were more likely to receive an LCS order within 1 day of their appointment (22.6% vs 16.4%; p=.010) and any time during follow-up (32.6% vs 24.1%; p=.002). Baseline LCS-LDCT completion was 51% higher in the intervention group, although the difference was not statistically significant (8.6% vs 5.7%; p=.078). Among the intervention group, 93 individuals (17%) viewed the video, generating 114 total views, and viewers watched an average of 79% of the video. Most views (82.5%) occurred through text-message delivery rather than QR codes. Conclusion: A brief, low-burden LCS informational video delivered immediately before a clinical encounter and integrated into existing workflows significantly increased LCS ordering and was associated with higher screening completion. Timely, scalable digital nudges may provide an effective strategy for improving LCS participation. Based on the observed effectiveness, feasibility, and efficiency of the intervention, KPCO incorporated the behavioral nudge into standard clinical care in February 2026.

14
Bayesian Borrowing of External Information in Clinical Trials: A Comparison of MAP, RMAP, and SAM Priors

Choi, L.; McNeer, E.; Beck, C. A.; Neul, J. L.

2026-08-31 pharmacology and therapeutics 10.64898/2026.08.26.26360843 medRxiv
Top 0.2%
1.4%
Show abstract

Bayesian borrowing of external information can improve trial efficiency, particularly in pediatric and rare disease settings where patient populations are limited, but may introduce bias and inflate the Type~I error rate when the trial differs from external studies. Recent U.S. Food and Drug Administration (FDA) draft Bayesian guidance emphasizes careful evaluation of external information, prior specification, and assessment of operating characteristics. This paper compares three meta-analytic-predictive (MAP)-based methods for Bayesian borrowing: the MAP prior, robust MAP (RMAP) prior, and self-adapting mixture (SAM) prior. An adaptive platform trial design in Rett syndrome is used as a case study. Simulation studies evaluate frequentist operating characteristics under varying prior--data conflict, between-study heterogeneity, treatment effects, and clinically significant differences (CSDs) for the SAM prior. The MAP prior achieved the greatest efficiency when external and current data were compatible but exhibited the largest bias under substantial prior--data conflict. The RMAP priors improved robustness through fixed robust-component weights, whereas the SAM prior adaptively adjusted borrowing and was less sensitive to prior--data conflict while retaining efficiency gains when the data were compatible. Although the CSD influenced the degree of adaptive borrowing, as reflected by effective sample size, it had only a modest impact on frequentist operating characteristics. Sensitivity analyses using a skeptical robust component yielded similar qualitative conclusions, while accentuating the differences between the MAP and RMAP priors. These findings provide guidance for evaluating and selecting MAP-based borrowing strategies before trial implementation, particularly in rare disease settings, consistent with current FDA recommendations.

15
Comparative Effectiveness of Single vs. Dual WhatsApp Reminders on No-shows: A Target Trial Emulation within the Public Health System of Buenos Aires, Argentina.

Esteban, S.; Quintana, G.; Sanchez, M.; Szmulewicz, A.

2026-08-19 health systems and quality improvement 10.64898/2026.08.17.26360609 medRxiv
Top 0.2%
1.4%
Show abstract

Background: Digital reminders reduce outpatient no-shows, but the optimal timing and frequency of messages remain unclear, particularly in Latin American public health systems. We emulated a target trial to evaluate the comparative effectiveness of four WhatsApp reminder strategies on appointment absenteeism and patient-initiated cancellations. Methods: We analyzed administrative and electronic health-record data from the public health system of the Autonomous City of Buenos Aires, Argentina (June 2023-May 2024). Eligible individuals had scheduled an in-person outpatient appointment in one of 15 prioritized specialties at least 75 hours in advance and had a mobile phone on record. We compared four strategies: (1) dual reminders at ~72 and ~24 hours before the appointment; (2) a single reminder at ~72 hours; (3) a single reminder at ~24 hours; and (4) no reminders. The primary outcome was the proportion of no-shows by the end of follow-up. Secondary outcomes were the cumulative incidence of patient-initiated cancellations overall, within 12 hours of the appointment, and followed by rebooking. We emulated the target trial using a cloning-censoring-weighting approach to estimate per-protocol controlled direct effects, with inverse-probability weights to address time-varying confounding and selection bias. Cumulative incidence of secondary outcomes was estimated using weighted Kaplan-Meier curves. Three pre-specified sensitivity analyses and standardized mean differences assessed robustness and covariate balance. Results: A total of 475,214 first eligible person-appointments were included; baseline no-show risk in the control arm was 34.6%. All three active strategies reduced no-shows compared with no reminders. The single 24-hour reminder produced the largest reduction (Risk Ratio [RR] 0.76, 95% CI 0.72, 0.81; Risk Difference [RD] -8.21 percentage points [pp], 95% CI -9.68, -6.54), followed by the dual-reminder strategy (RR 0.80, 95% CI 0.79,0.81; RD -7.05 pp, 95% CI -7.41, -6.71) and the single 72-hour reminder (RR 0.91, 95% CI 0.84,0.99; RD -3.16 pp, 95% CI -5.69, -0.49). All active strategies increased patient-initiated cancellations relative to control, with the dual-reminder strategy producing the largest increase. Sensitivity analyses preserved the qualitative ranking of strategies across all specifications. Conclusions: In this large target trial emulation, a single just-in-time WhatsApp reminder sent ~24 hours before the appointment was as effective as a dual-reminder schedule in preventing no-shows and superior to a distal 72-hour reminder alone. Adding a second, distal reminder provided no measurable benefit for attendance but substantially increased patient-initiated cancellations, which may be operationally valuable when active slot reallocation is a goal. These findings support timing, rather than frequency, as the primary lever of digital-reminder effectiveness, and favor the deployment of a single proximal reminder as the default strategy in resource-constrained outpatient settings.

16
Impact of Early Critical Care Pharmacist Involvement on Patient Outcomes in the Intensive Care Unit

Henry, K.; Smith, B. A.; Holden, D. N.; Smith, S. E.; Heavner, M. S.; Chen, Z.; Chen, X.; Devlin, J. W.; Murphy, D. J.; Martin, G. S.; Burden, M.; Murray, B.; Sikora, A.

2026-08-27 health systems and quality improvement 10.64898/2026.08.25.26361345 medRxiv
Top 0.2%
1.2%
Show abstract

Background: While critical care pharmacists (CCPs) are broadly associated with improvements in outcomes for critically ill patients, operationalizing staffing in the intensive care unit (ICU) requires further study. The purpose of this evaluation was to determine the relationship of a CCP on interprofessional rounds for weekday admissions of ICU patients on patient-centered outcomes. Methods: This post-hoc analysis of the Optimizing Pharmacist-Team Integration for ICU Patient Management (OPTIM) study included adults admitted to an ICU on a weekday in the multicenter observational study. The primary outcome was in-hospital mortality. The primary exposure was level of comprehensive medication management (CMM) during the first 24 hours of ICU stay. A secondary exposure was pharmacist-to-patient ratio. Multivariable generalized estimating equations (GEE) were used to estimate associations between mortality and patient, ICU, and institution variables. Fine-Gray sub-distribution hazards regression estimated hazard of discharge alive (HDA) from the ICU and hospital and hazard of extubation alive. Results: 21,835 patients met inclusion criteria, and 76.1% of patients had CMM delivered on interprofessional rounds. Patients who had no CMM on the first ICU day had an increased risk of mortality of 23% (Odds Ratio (OR) 1.23, 95% Confidence Interval (CI) 1.04-1.46, p=0.02) compared to those who received CMM on interprofessional rounds. Patients with no CMM also had decreased HDA from the ICU and hospital and decreased hazard of extubation alive. No difference was seen in any outcomes when comparing other levels of CMM (CMM delivered outside of interprofessional rounds or abbreviated CMM) compared to CMM delivered on rounds. Conclusions: Absence of pharmacist CMM on the first day of ICU stay for patients with weekday admission was associated with an increased risk of in-hospital mortality, but no difference was seen in other levels of CMM: this signal supports further investigation in prospective analysis.

17
Synthetic Longitudinal Tabular Data Generation via Copula

Cai, H.; Yu, W.; Lu, R.; Chattopadhyay, I.; Zhang, X.; Liu, J.

2026-08-07 bioinformatics 10.64898/2026.08.03.742474 medRxiv
Top 0.2%
1.1%
Show abstract

Synthetic data generation is increasingly used to enable data sharing and secondary analysis while protecting participant privacy, particularly for longitudinal tabular health data, where repeated measures per subject create within-subject dependence that most synthetic data methods are not designed to preserve. Existing generative methods, particularly generative adversarial network (GAN)-based approaches, can model complex distributions, but their estimated dependence structures are often difficult to interpret and their performance may be unstable or prone to overfitting in modestly sized datasets. Here we show that eCDF-copula, a statistically rooted approach using the empirical cumulative distribution function (eCDF) and copula modeling, preserves within- and between-visit dependence structure. To handle pervasive missing data, we propose a two-stage strategy combining multiple imputation with copula-based synthesis, enabling a variance decomposition that quantifies replication variability across methods. We benchmarked the proposed approach against four established methods on two longitudinal clinical datasets spanning markedly different sample sizes (n = 120 vs. n = 3, 612). eCDF-copula achieved resemblance and utility exceeding those of state-of-the-art synthetic data methods, while maintaining comparable privacy.

18
Estimating age-specific heterogeneity in SARS-CoV-2 transmission from prospective longitudinal studies: the importance of correcting for study design

Chervet, S.; Layan, M.; Boëlle, P.-Y.; Guedj, J.; van der Werf, S.; Kerneis, S.; Sermet-Gaudelus, I.; Cauchemez, S.; Opatowski, L.

2026-08-10 epidemiology 10.64898/2026.08.06.26358866 medRxiv
Top 0.2%
1.1%
Show abstract

Longitudinal household studies, combined with mathematical modeling, are widely used to characterize the drivers of respiratory pathogen transmission, including the effects of age and symptoms. In practice, household recruitment protocols vary across studies, potentially introducing biases into observed data. However, these biases are typically overlooked in statistical inference, and their impact on parameter estimates remains unknown. Here, we use synthetic household outbreak data simulated under different recruitment protocols to evaluate how recruiting through infected children affects estimates of age-specific infectiousness and susceptibility. We show that, under child-based recruitment, the standard likelihood, which accounts only for transmission dynamics, leads to underestimating child infectiousness and overestimating child susceptibility by more than 30%. We then propose a novel estimation framework that explicitly incorporates the household recruitment process into the likelihood and show that it substantially reduces these biases. Applying this new approach to a French household study conducted during the COVID-19 pandemic, we estimated that children under 6 had 49% lower infectiousness than teenagers and adults during the Alpha wave, whereas no difference was observed during the Omicron wave. This study demonstrates that ignoring recruitment protocols can bias key epidemiological parameter estimates and highlights the importance of accounting for study design.

19
Efficient Game-Theoretic Explanations for Tree-Based Ensembles via Owen Values

Koh, H.

2026-08-20 bioinformatics 10.64898/2026.08.12.744440 medRxiv
Top 0.2%
1.1%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWShapley-value-based explanations, notably SHAP (SHapley Additive exPlanations), have gained prominence as a principled game-theoretic framework for local explanations and global feature importance. While exact Shapley value computation is exponential in feature count, TreeExplainer exploits the recursive structure of decision trees to achieve polynomial-time computation for tree-based ensembles. In many scientific applications, however, features are naturally organized into a priori groups reflecting domain knowledge, requiring explanations both across and within groups. The Owen value extends the Shapley value through a two-stage allocation rule that incorporates group structure while preserving fairness properties; yet, efficient algorithms for its computation remain limited. In this paper, we propose exact and Monte Carlo algorithms for computing Owen values in tree-based ensembles by combining hierarchy-guided group aggregation with tree-aware dynamic programming. The exact algorithm computes Owen values without sampling under the path-dependent characteristic function, which approximates the conditional expectation, whereas the Monte Carlo algorithm provides a scalable approximation that is unbiased for any prespecified sampling budget and converges almost surely as the sampling budget increases. We also provide global importance measures and visualization tools for structured, multi-resolution explanations. The proposed algorithms and tools are collectively referred to as TreeOwen. Through simulation experiments, we demonstrate the numerical accuracy and substantial computational gains of TreeOwen. We illustrate its practical utility using immunotherapy metagenomic data, showing how microbial genera (groups) and species (features) contribute to patient recovery.

20
Likelihood-Based Inference and Model Selection for Stochastic Gene Expression in Probability-Generating-Function Space

Wang, Y.; Shu, Z.; McAuley, K. B.; Cao, Z.

2026-08-25 systems biology 10.64898/2026.08.24.746673 medRxiv
Top 0.2%
1.1%
Show abstract

Selecting stochastic gene-expression models from single-cell counts requires accurate parameter inference and efficient model selection. Likelihood methods in count space can be costly when full stationary count distributions are unavailable, whereas approximate methods may lose accuracy. Probability generating functions (PGFs) offer a compact analytical alternative, but existing PGF workflows are generally not likelihood based and therefore rely on computationally intensive cross-validation. We develop a likelihood-based PGF framework for both tasks. Correlated empirical PGF values are used to construct a Gaussian quasi-likelihood for parameter inference and PGF-based Bayesian information criterion (BIC) for model selection. We show that the empirical PGF is exactly unbiased and that the parameter estimator is consistent, converges at the inverse-square-root sample-size rate, and is first-order asymptotically unbiased. For large samples and a uniquely preferred model, PGF-BIC selects the same model as leave-one-out cross-validation in PGF space.