Biometrics
◐ Oxford University Press (OUP)
Preprints posted in the last 7 days, ranked by how well they match Biometrics's content profile, based on 23 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
McCready, T.; Thorpe, L.; Roy, B.; Renson, A.
Show abstract
Community-level estimates of healthcare utilization are essential for identifying inequities, allocating resources, and evaluating place-based interventions. However, in the United States, no single data source adequately captures healthcare utilization within geographically defined populations. Population-based surveys often lack sufficient geographic resolution, insurance claims represent only covered populations, and electronic health records are limited to care delivered within participating health systems. Increasingly, researchers combine these fragmented data sources, yet limited guidance exists for conducting valid population-based descriptive analyses using incomplete and overlapping data. We review the strengths and limitations of major data sources used to characterize community healthcare utilization and propose an approach for conducting population-based descriptive analyses using fragmented data. Rather than focusing on the limitations of individual data sources, our approach begins by explicitly defining the target population and the ideal observational study that would answer the research question. Available data sources are then conceptualized as incomplete or imperfect realizations of that ideal, providing a structured approach to (a) identifying sources of selection bias, missingness, and measurement error, (b) articulating required assumptions, and (c) selecting appropriate analytic strategies. We illustrate our approach using colorectal cancer screening utilization among adults residing in Brooklyn, New York during 2022. By shifting attention from individual data sources to the target community and the assumptions required for valid inference, this approach provides a practical approach for strengthening descriptive analyses of community healthcare utilization and informing place-based public health research, policy, and practice.
Liu, J. B.; Chen, Y.-J.; Edelen, M. O.; Pusic, A. L.; Martin, N. E.; Zeng, C.
Show abstract
Purpose: Nonresponse to routinely collected patient-reported outcome measures (PROMs) threatens the representativeness of aggregated data. We characterized patient-, provider-, and clinic-level factors associated with PROMIS Global-10 nonresponse in routine radiation oncology care. Methods: In this retrospective cohort study, all adults seen at five Mass General Brigham radiation oncology clinics over one year were included. The primary outcome was patient-level nonresponse, defined as never completing the portal-administered Global-10 versus completing it at least once. Using iterative mixed-effects logistic regression, we modeled patient-, provider-, and clinic-level factors. Results: Among 12,214 patients, 71 providers, and five clinics, patient- and appointment-level response rates were 35.4% and 10.9%, with patient-level response ranging nearly fivefold across clinics (12.8% to 66.2%). In Model 1, male sex, lower education, not working, and recent surgery had higher odds of nonresponse, and longer time since diagnosis lower odds. After provider- and clinic-level factors were added, patient sex, education, and employment became nonsignificant, whereas recent surgery (adjusted odds ratio [aOR] 1.97) and longer time since diagnosis (aOR 0.46 for >12 months) persisted. A provider's historical collection rate was protective but attenuated at the clinic level. There, a later program launch (aOR 0.29) and higher historical collection rate (aOR 0.79) correlated with lower nonresponse, whereas academic versus community setting did not. Conclusions: Nonresponse to routinely collected PROMs is a multilevel phenomenon driven substantially by clinic-level implementation factors, not patient characteristics alone. Because response rate is only a proxy for representativeness, PROMs programs and PRO-based performance measures should prioritize representative collection over volume.
Ng, S.-P.
Show abstract
The incidence rate ratio R is the standard measure for comparing event rates in clinical trials and epidemiology. In vaccine trials, the vaccine efficacy is VE = 1 - R. When events are rare, the two arm counts are Poisson. The estimator of R is heteroskedastic: its sampling variance changes with the data. So no fixed-width interval covers correctly everywhere. The usual log-Wald interval is undefined at zero events and covers poorly at small counts. Early vaccine and drug-safety readouts fall in exactly this regime. We show that a single reparameterization collapses this bivariate problem to an effective one-parameter family with a quadratic variance function, whose variance-stabilizing transformation is 2 arcsinh(sqrt(R)). The reduction yields a closed-form confidence interval for R. Its two leading errors, a curvature bias and the variability of the estimated scale, each admit a closed-form correction with no tuning constants. In a Monte Carlo study of our seven arcsinh variants and five competitors, the +Curve+Stu variant covers within 0.002 of the nominal 0.95 for about 50 control and 5 treatment events. Its width is on par with the best competitor. It avoids the conservatism and zero-count breakdown of log-Wald and MOVER. For moderate counts, we recommend this interval; for sparser data, our Bar-Lev and Enis count-shift variant is more robust. The result is a ready-to-use, closed-form interval for the low-count regime. We illustrate it on early Covid-19 vaccine-efficacy readouts and provide reference implementations in R and Python.
Frach, L.; Rijsdijk, F.; Hannigan, L. J.; Dudbridge, F.; Pingault, J.-B.
Show abstract
Polygenic scores are imperfect measures of the additive genetic effects of common genetic variants. The resulting measurement error biases estimates of quantities of interest in epidemiological analyses integrating polygenic scores. For example, how much of an exposure-outcome association is genetically confounded can be substantially underestimated when using polygenic scores alone. Here we present extensions to Gsens, a genetic sensitivity analysis, which aims to correct for such measurement error using both polygenic scores and heritability estimates. Gsens now allows for multiple exposures and estimates several quantities of interest, i.e. genetic confounding, adjusted residual association (net of genetic confounding), genetic overlap and environmentally mediated genetic effects. We present derivations and simulations showing how Gsens accounts for measurement error in the polygenic score; we also show how estimation may be affected by misspecifications of the causal structure between exposures. Applying Gsens in the Norwegian Mother, Father and Child Cohort Study (MoBa), we uncover, among other results, substantial genetic confounding in the associations between multiple known risk factors for attention deficit hyperactivity disorder (ADHD), such as low birth weight and temperament, and measures of ADHD in childhood. The updated Gsens R package offers multiple options, including for missing data handling and customisable syntax. Our extended version of Gsens is applicable to a broad range of substantive questions in multiple disciplines.
Hamilton, F. W.; Ong, S. Y.; Swets, M.; Russell, C. D.; Underwood, J.
Show abstract
Background Staphylococcus aureus bacteraemia (SAB) is clinically heterogeneous. Potential heterogeneous treatment effects (HTE) have recently been identified through analysis of patient subgroups, identified using routine clinical variables.However, the impact of misclassifying patients into these groups is unclear, and practical strategies to improve HTE detection remain uncertain. Methods We performed a simulation study using published data from selected randomised trials and observational studies in SAB. We assessed the impact of varying classification accuracy (70%-100%) on i) power, ii) type I error, and iii) bias in post-hoc analyses of HTE. We then evaluated two strategies to improve performance: enrichment designs, in which only patients predicted to belong to a target subgroup are randomised, and the use of ordinal rather than binary outcomes. Results Even with perfect classification, post-hoc detection of heterogeneous treatment effects remained highly conditional on subgroup prevalence, baseline mortality, and effect size. One subgroup was detectable at moderate sample sizes; however, power was inadequate for all other subgroups even with sample sizes of 20,000. Decreasing classification accuracy reduced power, increased type I error, and introduced bias. Enrichment marginally improved power. Ordinal outcomes substantially improved performance when they matched the treatment-effect structure, but were worse when they did not. Conclusions Detecting HTE in SAB is challenging, but not uniformly infeasible. Feasibility depends on the interaction between subgroup frequency, baseline risk, classifier performance, and outcome choice. To advance stratified medicine in SAB, research should prioritize robust classifiers, outcome measures matched to the expected mechanism and pattern of treatment effect, and trial designs that acknowledge uncertainty in subgroup prevalence and treatment-effect structure.
Velasco Pardo, V.; Daines, L.; Katikireddi, S. V.; Ritchie, L.; Robertson, C.; Simpson, C. R.; McCowan, C.; Swallow, B.
Show abstract
Background During the COVID-19 pandemic, public health agencies used near real-time observational data to answer questions regarding vaccine effectiveness. However, traditional observational methods do not allow conclusions regarding counterfactual scenarios to be drawn from clinical data. Counterfactuals, which are outcomes that would have occurred under alternative interventions, can be used to formally assess the causal effects of public health interventions on health outcomes while accounting for the effects of confounding. Ideally individual patient data is used for the development of counterfactuals. Low-fidelity synthetic data may be useful for advancing methodological development where governance and privacy constraints prohibit access to sensitive personal data. Methods We simulated synthetic datasets based on the EAVE-II COVID-19 platform which has been limited to use for surveillance purposes. EAVE-II includes almost all resident people in Scotland registered with qualified general medical practitioners. Patient characteristics were simulated to reflect the known distribution of the Scottish population, accounting for dependencies between variables. Each synthetic dataset was encoded to different realistic scenarios for EAVEII 'ground truth' vaccine rollout and effectiveness results, explicitly stating the causal and confounding mechanisms, using a statistically sound method based on marginal structural models. Synthetic datasets of 100,000 individuals were then generated across five confounding scenarios and five severe outcome types. Results In scenarios with weak confounding, both unweighted and inverse probability of treatment weighted (IPTW) logistic regression recovered the true causal parameters. As confounding strength increased, only weighted models recovered the true mechanism. Conclusions Low-fidelity synthetic datasets simulated from EAVE-II data analysts to build and test causal inference pipelines, develop novel analysis pipelines, and train new researchers while awaiting access to real data. We showed how to generate synthetic datasets from a marginal structural model under different confounding scenarios.
Hsu, C.-Y.; Liu, Q.; Shyr, Y.
Show abstract
As machine learning and artificial intelligence systems are increasingly used in healthcare, rigorous evaluation of their classification performance has become critical. The F1 and F{beta} scores are widely adopted metrics for assessing performance in imbalanced biomedical data. Recently, we introduced psF1, a unified statistical framework for inference and study design for single and comparative F1 and F{beta} scores under the assumption of independent classifiers. In practice, however, benchmarking two classifiers on the same dataset creates a correlated paired setting. Ignoring this intrinsic dependency leads to overestimation of the standard error and a substantial loss of statistical power. To address this, we develop psF1pair, an advanced framework for statistical inference and power analysis that explicitly accounts for correlations between classifier pairs. Extensive simulation studies demonstrate the performance of psF1pair, and its utility is further illustrated through application to a real-world imaging classification system. As expected, higher correlation between classifiers yields narrower confidence intervals and enhanced statistical power. A freely available R package is provided to facilitate implementation, supporting accurate evaluation and study design for predictive and classification models in biomedical research.
Gusinow, R.; Morgan, A. S.; Canziani, L. M.; Zeitlin, J.; Kim, M.; Gentilotti, E.; Ghosn, J.; Florence, A.-M.; Tami, A.; Toschi, A.; Palacios-Baena, Z. R.; Tacconelli, E.; Hasenauer, J.
Show abstract
Causal effect estimates can often be biased in clinical and epidemiological studies as patient cohorts frequently exhibit substantial covariate imbalances between treated and control groups, often amplified in multicentre studies due to heterogeneous recruitment, clinical practice, and case mix. Covariate balancing methods are therefore essential for valid causal inference. However, their application becomes challenging when data are distributed across cohorts and cannot be pooled because of privacy, legal, or institutional constraints, leaving a gap in practical methods for causal effect estimation in federated and imbalanced clinical data settings. We develop a privacy-preserving framework for covariate balancing and causal effect estimation across distributed data providers, combining federated aggregation with differential privacy to enable propensity score subclassification and matching without sharing individual-level records. Matching relies on non-disclosive quantities and differentially private distance evaluation, and the resulting matched subsets remain local to each server. Balance can be assessed through federated diagnostics and privacy-preserving visualisations, and we provide secure estimators for average treatment effects with associated uncertainty quantification. We implement this framework in the DataSHIELD federated analysis platform via 2 R packages. In simulations, we demonstrate agreement between federated and centralised analyses in the absence of privacy noise and quantify the bias--variance trade-offs induced by differential privacy. We illustrate applicability in two multinational settings-a Long COVID cohort and very preterm birth cohorts-showing that the approach enables practical causal analyses under real-world data protection constraints. The DataSHIELD packages are available on Github. Additional methodological details are provided in the Supplementary Material.
Erly, B.; Raja, S.
Show abstract
Background. Compounded tirzepatide is prescribed at scale as a cheaper substitute for branded Mounjaro and Zepbound, yet the cost case is almost always built by setting one compounded price against one branded list price. That framing ignores the question that actually decides the answer: cheaper than which branded price the patient can reach. Branded tirzepatide is now sold at sharply different tiers, namely insurance copay (often $25-$150/month), LillyDirect Self Pay ($299-$449/month), and retail cash price ($1,000-$1,200/month). Whether compounded saves money turns entirely on which of these a given patient faces. A second open question is whether the two formulations even produce comparable effectiveness, since observed differences may reflect selection on insurance, baseline characteristics, and adherence rather than the drug. Methods. We conducted a retrospective cohort study of tirzepatide users in the Mochi Health telehealth program, classified by formulation from their refills as branded-only (Mounjaro/Zepbound; 6,238), compounded-only (71,683), or switchers (4,996); switchers were excluded from the formulation contrast. Among single-formulation patients with a documented six-month weight observation, the analytic cohort was 7,271 (869 branded, 6,402 compounded). The primary outcome was six-month percent body weight loss; the secondary outcome was >=10% response. We used 1:1 nearest-neighbor propensity-score matching (0.25 SD caliper) on baseline covariates only - age, sex, baseline BMI, baseline weight, comorbid diabetes, hypertension, dyslipidemia, prior bariatric surgery, and self-reported insurance coverage - deliberately excluding post-treatment variables such as adherence and time in program, which are mediators of the formulation effect. We pre-specified an equivalence margin of +/-2 percentage points on mean loss and tested equivalence with two one-sided tests (TOST). A directed acyclic graph (DAG) makes the identifying assumptions explicit; metformin use could not be reliably ascertained and is treated as an unmeasured confounder. The cost comparison reports the savings or premium of compounded versus branded under five branded price scenarios: retail list, LillyDirect Self Pay (two dose tiers), and insurance copay (typical and low end). It is a cost comparison (cost-minimization under demonstrated similar effectiveness), not a formal cost-effectiveness analysis: we computed no ICER, QALY, or discounting. Results. Branded and compounded patients had similar outcomes even before adjustment (mean loss 11.7% vs 11.5%; >=10% response 60.9% vs 58.8%). The largest baseline difference between the groups was insurance coverage (branded patients far more likely insured; standardized mean difference 0.67), which matching balanced to 0.01. After 1:1 matching (718 pairs, all |SMD| < 0.04), mean loss was 11.4% vs 11.4% (difference +0.08 pp, 95% CI -0.70 to +0.80) and >=10% response 59.3% vs 57.2% (difference +2.1 pp, 95% CI -3.1 to +7.1). The two formulations were statistically equivalent within the pre-specified +/-2 pp margin (TOST p < 0.001). Cost depends on the branded scenario: compounded saves $6,000 over six months versus retail list price, $1,494 versus LillyDirect maintenance-dose (5-15 mg) Self Pay, and $594 over a low-dose (2.5 mg) LillyDirect prescription, while it costs $300 more than branded under a typical insurance copay ($150/month) and is more expensive still at lower copays (savings turn negative below $200/month). Conclusions. Branded and compounded tirzepatide were statistically equivalent in six-month effectiveness within a pre-specified +/-2 pp margin, so the choice between them is essentially a cost decision - and that cost advantage is real but conditional on the branded price the patient can access. It is large against retail list price and shrinks to zero or reverses against LillyDirect Self Pay or a low insurance copay. Whether compounded is the lower-cost choice for an individual patient is, therefore, a question about which price tier that patient faces.
Williams, J.; Mencer, N.; Mak, W. Y.; Dalle Luche, G.; Dundovic, S.
Show abstract
Background Hypertension is a major modifiable risk factor for atrial fibrillation (AF), yet blood pressure (BP) control remains suboptimal in older U.S. adults. Objectives This study evaluated how improve systolic BP (SBP) control could affect incident AF, downstream AF ablation demand, Medicare savings, and hospital revenue. Methods A population-based modelling framework was developed to estimate mortality and incident AF hazards across SBP strata: <120, 120-139, 140-159, and ?160 mm/Hg. AF incidence in the SBP <120 mmHg group was set at 2.2 per 1,000 person-year, with hazard ratios of 1.17, 1.42 and 1.64 applied to higher SBP strata. We assumed 25% of incident AF patients would undergo ablation, with a 7.2% complication rate. AF prevalence was projected to increase by 4.6% annually over 10 years. Medicare savings and hospital revenue foregone were estimated under varying procedure cost and contribution-margin assumptions. Results Higher SBP was associated with greater hazards of death and incident AF. Improved SBP control reduced projected AF incidence and ablation demand. Over 10 years, cumulative Medicare savings were projected at $8.7B-$10.9B across the full modelled population. However, reduced ablation volume translated into hospital revenue foregone, ranging from $75M to $377M in the first year, and approximately $1.03B-$5.2B cumulatively over 10 years. Conclusions Improved SBP control may reduce AF incidence, prevent avoidable invasive ablation procedures, relieve pressure on surgical waitlists, and generate substantial Medicare savings. However, these benefits may reduce hospital procedural revenue, highlighting a misalignment between prevention-oriented care and fee-for-service reimbursement incentives.
Erly, B.; Raja, S.
Show abstract
Background. When a GLP-1 patient stops responding, should the clinician push the dose or change the drug? Observational answers conflate two distinct sources of confounding. Most early-week "escalations" in real-world data are FDA-mandated titration steps rather than deliberate clinical decisions, and patients who deviate do so for reasons we cannot observe. Semaglutide and tirzepatide are also not equivalent milligram-for-milligram, so naive class-switch comparisons mix mechanism and dose. We resolve both by restricting to post-titration patients and reclassifying treatments under the Whitley 2023 dose-equivalence framework. Methods. From 68,969 telehealth GLP-1 patients we built a post-titration cohort. Each patient's index time is the day they completed at least four weeks at therapeutic dose (Whitley tier 3 or higher: semaglutide 1.0 mg or tirzepatide 5 mg). Confirmed slow response is less than 5% total weight loss at the index, consistent with FDA weight-management drug-development guidance and AACE/ACE criteria. We compared four post-index strategies against continuing the current regimen: within-class dose escalation, equipotent class switch (a Whitley tier change of 1 or fewer), and class switch with potency increase. Direction-specific analyses split switches into semaglutide-to-tirzepatide and tirzepatide-to-semaglutide arms. Outcomes were percent weight loss at 12 and 24 weeks post-index. We estimated effects six ways: propensity-score matching; IPTW with linear and gradient-boosted propensities; the g-formula with linear and gradient-boosted outcome models; and AIPW, the doubly-robust estimator we use as the tiebreaker. Two-layer inverse probability of censoring weighting addressed strategy adherence and outcome ascertainment. We computed E-values, ran a negative-control specification, and stratified by tolerability. Results. The post-titration cohort comprised 24,876 confirmed slow responders. Within-class dose escalation produced a small consistent benefit at 24 weeks: AIPW +0.64 pp (95% CI +0.16 to +1.12), with five non-AIPW estimators ranging +0.47 to +0.76 pp. The continue arm itself lost an additional 8.07 pp over the same window (96% continued to lose), so escalation is a marginal addition to a substantial natural slope, not a rescue. Equipotent class switching from semaglutide to tirzepatide was inconclusive: linear and matching estimators ranged +1.26 to +1.80 pp, but AIPW was -0.33 pp (95% CI -1.27 to +0.60) with only 90 treated patients and limited propensity-score overlap. Class switch with simultaneous potency increase (sema to tirz) gave AIPW +0.65 pp at 12 weeks (95% CI +0.37 to +0.92, n = 80). A negative-control specification yielded ATE -0.13 pp, indicating the pipeline did not generate spurious signal. A held-out-fold prognostic-threshold sensitivity gave a null effect (+0.04 pp), correcting an earlier circular +1.06 pp estimate. Conclusions. Among confirmed post-titration slow responders, within-class dose escalation adds approximately 0.6 percentage points at 24 weeks on top of an 8 percentage point natural slope, consistently across six estimators including doubly-robust inference. This headline effect is small and not robust to modest unmeasured confounding (E-value 1.27) or to MNAR-style outcome attrition (tipping point delta approximately 1.2 pp), so it should be read as hypothesis-generating rather than practice-changing. Class-switching evidence is inconclusive; linear-estimator results suggesting benefit did not survive doubly-robust estimation in small treated samples with limited propensity overlap. The dose-ladder framework, with phase-specific evidence grading, is hypothesis-generating and insufficient on its own to change practice.
Weerasinghe, C.; Osowicki, J.; Simpson, J. A.; Crocker-Buque, T.; McCarthy, J.; Williams, E.; Price, D. J.
Show abstract
Controlled human infection models (CHIMs) are increasingly used in infectious disease research to study pathogen dynamics and evaluate interventions under controlled conditions. However, these studies are resource-intensive and involve ethical and safety constraints, making efficient study design critical. Dose-finding is a key early component in CHIMs, where the aim is to identify a challenge dose that achieves a target infection probability. Traditional rule-based designs are commonly used but can be inefficient, motivating the use of model-based adaptive approaches such as the Bayesian Continual Reassessment Method (CRM). Although CRM has been extensively studied and widely adopted in Phase I oncology trials for identifying the maximum tolerated dose of therapeutics, its application in CHIM settings remains limited, particularly when the endpoint of interest is infection. This tutorial provides step-by-step guidance for implementing a Bayesian CRM in dose-finding CHIMs, using an oropharyngeal Neisseria gonorrhoeae challenge as a motivating case study. The framework outlines key design components, including dose-grid specification, dose-response model, prior elicitation, Bayesian updating, decision rules, and stopping criteria, with particular emphasis on a clinically interpretable parameterisation. Trial operating characteristics are evaluated through simulation studies under multiple dose-response scenarios and prior-predictive analyses, and compared with a commonly used '3+3' type rule-based design. This work highlights the advantages of Bayesian model-based designs for dose-finding in CHIMs over classic rule-based designs and provides a structured, reproducible framework for implementing CRM, supporting their application in future CHIM studies.
Wang, Z.; liu, y.
Show abstract
Background Primary care in China lacks structured mental-health assessment, and the machine-learning models that could support such screening are typically developed on heavily selected samples. Cumulative inclusion and exclusion criteria, though usually treated as neutral data-cleaning steps, can create heterogeneity in predictive reliability among retained participants. Using the China Health and Retirement Longitudinal Study (CHARLS) 2011 baseline, we quantified how selection funnels distort epidemiological associations and inflate machine-learning metrics, and tested selective prediction as mitigation. Methods Using the CHARLS 2011 baseline with temporal external validation in CHARLS-2018, we built a four-level selection funnel (L0-L3), evaluated five classifiers with nested cross-validation and SMOTE, and compared model-embedded uncertainty with a decoupled predictor-selector framework; XGBoost cross-validation residuals drove risk stratification and classification and regression tree (CART) rules. Results Sample sizes fell from L0 n=17,705 to L3 n=4,256 (24.0%). The cancer-depression odds ratio attenuated from 1.78 (95% CI 1.32-2.41) to 1.39 (0.74-2.63), losing significance. AUC rose with selection but not after multiple-comparison correction, whereas calibration error increased for four of five models. Model-embedded uncertainty succeeded only for XGBoost; with the decoupled XGBoost residual selector, all five models achieved selective prediction at approximately 20% coverage (test AUC 0.90, 95% CI 0.85-0.95), abstaining on approximately 80% of cases for individual safety. Risk stratification was stable (residual Spearman correlations >0.95; multi-seed Jaccard 0.88), and CART rules used self-rated health, education, pain, and marital status. Conclusions The findings support a deployable primary-care triage pathway: a four-variable rule identifies patients suitable for algorithm-assisted scoring (approximately 20% coverage) and routes the remainder to human evaluation. Methodologically, cumulative selection bias produces a dual distortion: epidemiological associations are compressed and machine-learning metrics inflated. Selective prediction is limited mainly by uncertainty-indicator design. Performance metrics should be reported with selection level, coverage, and calibration trajectory. Decoupled selective prediction with CART rule extraction provides an actionable framework for quality-controlled, tiered-care deployment. Keywords: selective prediction, selection bias, CHARLS, depression, predictor-selector decoupling, uncertainty quantification, classification and regression tree, triage, clinical decision support, health management.
van Boven, M.; Bootsma, M. C.
Show abstract
Stochastic epidemic models are a cornerstone of infectious disease epidemiology and are often used to study intervention scenarios. However, large run-to-run variability can make intervention effects difficult to estimate precisely. We revisit the epidemic Sellke construction, which assigns each individual an infection threshold for the cumulative infection hazard such that, conditional on the thresholds, the epidemic trajectory becomes deterministic. This enables coupling of simulations with and without an intervention, yielding low-variance effect estimates even when outcomes such as final size or peak incidence vary widely between runs. We develop an exact, event-driven implementation that maintains infection and recovery events in priority queues. Cumulative infection-hazard updates require O(log N) time per event, yielding overall complexity O(Elog N) for E events in a population of size N. The implementation achieves computational performance comparable to the classical Gillespie algorithm while naturally accommodating non-Markovian infectious periods and complex infectiousness profiles. We illustrate the approach using distance-dependent spread of avian influenza between poultry farms in the Netherlands and a multilayer population with households, schools, and workplaces. In both examples, coupling enables efficient within-run comparisons of intervention scenarios across stochastic realisations.
Suresh, J.
Show abstract
Malaria subnational tailoring is often a population-level allocation problem: which interventions should be prioritized, at what coverage, and under what budget and uncertainty assumptions? We present DELENDA, a differentiable compartmental model of Plasmodium falciparum transmission designed for posterior calibration and intervention-mix optimization. We fit a NUTS posterior jointly to age-stratified prevalence and clinical-incidence data from five sub-Saharan African sites plus three pre-intervention Garki Project villages, spanning a broad entomological inoculation rate (EIR) range. DELENDA is implemented in JAX, which makes the full simulation differentiable. This enables efficient Bayesian inference and continuous constrained optimization over intervention coverage. We apply the framework to an illustrative decision problem: a highly seasonal transmission setting where coverage is optimized for ITNs, SMC, IRS, and pediatric malaria vaccination across EIR, budget, objective, and uncertainty grids. Three findings are decision-relevant. First, intervention rankings are more robust than projected impact: posterior, vector-biology, and intervention-efficacy uncertainty change optimized coverage modestly but substantially widen the distribution of cases averted. Second, the objective matters: under-five optimization brings child-targeted SMC and vaccination in earlier, whereas all-age optimization delays vaccination and favors broader population protection through IRS. Third, cost uncertainty is mainly a constraint-side problem: expected-cost optima have material budget-overrun probability, while tail-risk budget rules sharply reduce overrun risk at the cost of lower effective coverage and fewer expected cases averted. DELENDA therefore demonstrates an uncertainty-first approach to subnational tailoring: differentiable model structure exposes the biological parameter space to posterior calibration and carries biological and operational uncertainty into constrained decision optimization, tasks that are difficult with the non-differentiable models currently central to SNT workflows.
Cheng, C.
Show abstract
Genome-scale Perturb-seq screens prioritize candidate targets by the strength of a perturbations transcriptional effect. Effect strength does not answer a prior measurement question: is the readout dependable? A large effect estimated from a single guide, a single donor, or a pseudobulk of few cells need not survive replication, and for target prioritization each false lead costs a validation experiment. We treat each perturbation effect as a measurement in a crossed Target x Guide x Donor x Condition design and apply generalizability theory (Brennan, 2001; Cronbach et al., 1972) to separate the dependable part of an effect from facet-specific idiosyncrasy. Guides and donors enter as random facets; condition enters as a fixed facet and is analyzed within its levels. For each target we report a dependability profile over the facets and a joint generalizability coefficient over the two random facets, and we re-rank targets by effect magnitude weighted by that coefficient. On the released screen (Zhu et al., 2025), removing the measurement-error floor estimated from the non-targeting controls raises the number of genes with a dependable target-signal share above .10 from 40 to 7,674. Analyzed within activation states, dependability recovers the T-cell-receptor signaling module as reliably measurable only in activated cells, without recourse to gene annotation. A design study indicates that reliability is limited by the number of guides rather than the number of donors, so a future screen should add guides. Every methodological decision was recorded and adversarially reviewed, and all results regenerate from the released summary statistics.
Taliun, D.; Gagliano Taliun, S. A.
Show abstract
As population-scale whole-genome sequencing datasets continue to expand, they enable genetic association studies beyond single-nucleotide variants to more complex forms of genetic variation, including classical human leukocyte antigen (HLA) alleles. The HLA region comprises nine highly polymorphic classical HLA genes in extensive linkage disequilibrium that are associated with numerous autoimmune and infectious diseases. However, unlike genome-wide association studies of single-nucleotide variants, there is no general guidance for controlling the multiple-testing burden in HLA allele association analyses. Here, we systematically evaluated the effective number of independent HLA allele tests using sequencing data from diverse genetic ancestries, analytical derivation and simulations. We show that the multiple-testing burden depends on genetic ancestry, allele frequency, and the phenotype model, but remains remarkably stable across minor allele count thresholds, corresponding to approximately 60-70% of the total number of tested HLA alleles. Simulations further demonstrate that the effective number of tests can exceed 90% under realistic disease models. Analyses of 4-field HLA alleles from long-read sequencing showed that higher typing resolution increases the number of alleles but preserves the underlying correlation structure and scales the effective number of independent tests proportionally. Our results provide practical guidance for HLA association studies and support Bonferroni correction based on the total number of tested HLA alleles as a simple and robust approximation when permutation-based approaches are impractical.
Samuels, T. H.; Forrest-Hammond, R.; Stockford, C.; Harris, S. K.; Eyre, D. W.; Gupta, R. K.; Noursadeghi, M.
Show abstract
Background: Bacteraemia is associated with poor outcomes but the diagnostic gold standard, peripheral blood culture, takes up to 24 hours to become clinically actionable, hampering early management decisions in suspected infection. Single predictors and existing sepsis risk scores discriminate poorly, and few multivariable bacteraemia models have been adequately validated in UK populations. Methods: We developed a logistic regression model, using backwards AIC based selection of predefined candidate predictors routinely available within hours of hospital attendance, in a retrospective cohort of 33,874 hospital encounters at University College London Hospitals (UCLH) between 2019 and 2024. Continuous predictors were modelled using restricted cubic splines and missing data handled using multiple imputation. Model performance was assessed via internal external cross validation and prediction instability analysis, before temporal validation in held-out 2024 UCLH data and external validation in 53,669 hospital encounters from the Infections in Oxfordshire Research Database (IORD). Results: Bacteraemia occurred in 5.2% of UCLH and 8.9% of IORD encounters, respectively. Twenty predictors were retained, spanning demographics, comorbidities, vital signs and blood tests. Discrimination was stable across development time periods (pooled c-statistic 0.82, 95%CI 0.81 to 0.84) and was maintained in temporal (0.83, 0.79 to 0.87) and external validation (0.83, 0.82 to 0.83), with excellent calibration in external validation (calibration slope 1.08 (1.05 to 1.11); calibration-in-the-large 0.01 (-0.02 to 0.04)). The model outperformed single predictors, established risk scores, and a reconstructed comparator model, and showed superior net benefit in decision curve analysis. Performance was consistent across age, sex, ethnicity and socioeconomic subgroups but degraded when blood cultures were sampled more than six hours after attendance and varied by likely infection site. Conclusions: This model accurately predicts bacteraemia using routinely collected data available within hours of hospital attendance, with performance maintained in a large, independent external validation cohort. It offers a generalisable, clinically interpretable tool to support early decision-making in suspected infection, pending further work to establish optimal implementation thresholds.
Amiri, S.; Afshar, P.; Rohban, M. H.
Show abstract
Objectives. Radiomics pipelines extract hundreds of quantitative features that are widely known to be redundant, but the structure of this redundancy is usually treated as a per-dataset nuisance to be pruned away. We tested the alternative hypothesis that a substantial number of feature-feature correlations are universal: they persist across patients and across anatomically distinct structures because they reflect shared mathematical and image-statistical properties of how the image is summarised, rather than properties of the tissue being imaged. Materials and Methods. We re-analysed the publicly available Radiomics Atlas Dataset of normal Abdominal and Pelvic CT (RADAPT), restricting the analysis to the 526 non-contrast-enhanced examinations of the 531-subject atlas and to the 107 original (non-filtered) PyRadiomics features. The 53 segmented structures were grouped into four broad anatomical categories -- bones, muscles, vessels, and parenchymal organs. RADAPT is distributed as one Excel file per structure, with patients as rows and features as columns. Within each structure file we z-score-normalised every feature across patients, computed the absolute Spearman correlation matrix, and retained edges with |{rho}| [≥] {tau} for {tau} in {0.70, 0.80, 0.90}. We then intersected the edge sets across all structure files to obtain a "universal" correlation graph, in which an edge survives only if it exceeds the threshold in every structure (each estimated across the full patient sample). Stable feature communities were defined as the maximal cliques of this graph. Robustness to patient sampling was tested by repeating the entire pipeline on five independent random splits of each file into two patient halves (10 sub-cohorts per threshold), and the implementation was independently reproduced in R. Results. Despite the strictness of the global-intersection criterion, 34, 24, and 14 stable feature communities survived at {tau} = 0.70, 0.80, and 0.90 respectively, with the largest cliques containing six members at {tau} = 0.70 and {tau} = 0.80 and five members at {tau} = 0.90. The community structure was clearly interpretable: separate cliques captured (i) variance-like intensity dispersion, (ii) long-run / low-frequency (coarse) texture, (iii) high gray-level texture, (iv) low gray-level texture, (v) volume and surface shape, and (vi) local-homogeneity and energy/entropy duals. On random-half resampling the exact-match recovery rate of these communities was 81.5 %, 86.7 %, and 80.7 % across the three thresholds; departures from exact recovery were almost always a single boundary feature added or dropped, consistent with finite-sample fluctuation of near-threshold edges rather than structural instability. The R re-implementation reproduced the Python results exactly. Conclusion. A substantial portion of radiomics feature collinearity is universal across patients and tissues. We distinguish two layers within it: trivial near-algebraic duals that are universal by construction, and non-trivial cross-matrix-family communities that are the genuine empirical finding. Together they provide an interpretable, definition-grounded basis for aggressive dimensionality reduction, for retrospectively reconciling apparently different feature selections in the literature, and for moving radiomics pipelines toward organ-agnostic, more reproducible models. Clinical relevance statement. Selecting a single representative feature from each universal community shrinks the original-feature space by roughly an order of magnitude without sacrificing biologically distinct information. For example, the five variance-family members (first-order Variance, GLCM SumSquares, GLCM ClusterTendency, GLDM and GLRLM GrayLevelVariance) can be replaced by a single representative, removing redundant degrees of freedom that would otherwise inflate model variance; and labelling each retained feature by its community lets two studies that selected different variance-family names be recognised as having found the same signal, simplifying model development and improving cross-cohort generalisability in clinical CT workflows.
Rahman, M. M.; Guha Niyogi, P.
Show abstract
The apnea-hypopnea index (AHI), the conventional metric of obstructive sleep apnea (OSA) severity, is typically studied using scalar summaries of sleep architecture, such as the total time spent in each sleep stage. Although clinically interpretable, these summaries fail to capture the temporal organization of overnight sleep-stage sequences and may obscure stage-specific associations with OSA severity. Modeling the complete sleep-stage trajectory provides substantially richer temporal information; however, because total sleep duration varies across individuals, sleep-stage trajectories are observed over subject-specific domains, limiting the applicability of conventional functional regression methods that assume a common observation interval. We therefore applied Variable-Domain Functional Regression (VDFR) to overnight polysomnographic data from the APPLES study (n= 1,103), treating the epoch-by-epoch sleep-stage sequence as a continuous, variable-length functional predictor of AHI. We compared three levels of sleep-stage granularity: five stages (Wakefulness, N1, N2, N3, REM), three stages (Wakefulness, Non-REM, REM), and binary staging (Wakefulness vs. Sleep). Functional sleep-stage terms were significant across all staging granularities and model structures (all p-values [≤]0.001). Wake, N1, and N2 were positively associated with AHI, whereas N3 and REM were negatively associated, with REM exhibiting the strongest association. These effects were attenuated under coarser staging representations, highlighting the importance of preserving fine-grained sleep architecture. To our knowledge, this is the first application of VDFR to overnight polysomnographic data in OSA, showing that accommodating subject-specific sleep durations enables the identification of stage-specific temporal associations with AHI severity that are attenuated or obscured by coarser staging and conventional scalar analyses.