Biostatistics
◐ Oxford University Press (OUP)
All preprints, ranked by how well they match Biostatistics's content profile, based on 24 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Hatami, F.; Perrakis, K.; Cooper-Knock, J.; Mukherjee, S.; Dondelinger, F.
Show abstract
SO_SCPLOWUMMARYC_SCPLOWLarge-scale longitudinal data are often heterogeneous, spanning latent subgroups such as disease subtypes. In this paper, we present an approach called longitudinal joint cluster regression (LJCR) for penalized mixed modelling in the latent group setting. LJCR captures latent group structure via a mixture model that includes both the multivariate distribution of the covariates and a regression model for the response. The longitudinal dynamics of each individual are modeled using a random effect intercept and slope model. Inference is done via a profile likelihood approach that can handle high-dimensional covariates via ridge penalization. LJCR is motivated by questions in neurodegenerative disease research, where latent subgroups may reflect heterogeneity with respect to disease presentation, progression and diverse subject-specific factors. We study the performance of LJCR in the context of two longitudinal datasets: a simulation study and a study of amyotrophic lateral sclerosis (ALS). LJCR allows prediction of progression as well as identification of subgroups and subgroup-specific model parameters.
Hutchings, G.; Samartsidis, P.; Donnay, C.; Gaetano, L.; Fisher, E.; Nichols, T. E.; Holmes, C.; Häring, D. A.; Ganjgahi, H.
Show abstract
Probabilistic latent variable models are a powerful tool for uncovering structure in high-dimensional datasets, particularly in biomedical applications. The increasing availability of large-scale epidemiological studies, such as the UK Biobank, poses important modelling challenges, including mixed data types, high dimensionality, and structured missingness. Existing approaches address some of these issues, but few provide a unified and scalable framework for handling them simultaneously. Here, we propose a scalable Bayesian factor analysis framework designed to address these challenges. Our method combines a semi-parametric Gaussian copula model with a continuous spike-and-slab prior to induce sparse and interpretable factor loadings. The number of latent dimensions is learned nonparametrically from the data using an Indian buffet process prior. For model fitting, we develop an expectation-maximisation algorithm that naturally accommodates missing data. We validate the proposed method through comprehensive simulation studies. In addition, we showcase the proposed model using the Novartis-Oxford Multiple Sclerosis dataset in two ways. First, we identify latent dimensions shared across MS clinical and neuroimaging variables, characterising disease structure while demonstrating the models ability to handle multiple data types and structured missingness. Second, we use the model for dimensionality reduction of structural MRI data, extracting features for downstream analysis that go beyond traditional whole-brain summary statistics. In those applications, our method identifies sparse latent structures and provides insights beyond those obtained from traditional approaches.
Ahuja, Y.; Liang, L.; Huang, S.; Cai, T.
Show abstract
Leveraging large-scale electronic health record (EHR) data to estimate survival curves for clinical events can enable more powerful risk estimation and comparative effectiveness research. However, use of EHR data is hindered by a lack of direct event times observations. Occurrence times of relevant diagnostic codes or target disease mentions in clinical notes are at best a good approximation of the true disease onset time. On the other hand, extracting precise information on the exact event time requires laborious manual chart review and is sometimes altogether infeasible due to a lack of detailed documentation. Current status labels - binary indicators of phenotype status during follow up - are significantly more efficient and feasible to compile, enabling more precise survival curve estimation given limited resources. Existing survival analysis methods using current status labels focus almost entirely on supervised estimation, and naive incorporation of unlabeled data into these methods may lead to biased results. In this paper we propose Semi-supervised Calibration of Risk with Noisy Event Times (SCORNET), which yields a consistent and efficient survival curve estimator by leveraging a small size of current status labels and a large size of imperfect surrogate features. In addition to providing theoretical justification of SCORNET, we demonstrate in both simulation and real-world EHR settings that SCORNET achieves efficiency akin to the parametric Weibull regression model, while also exhibiting non-parametric flexibility and relatively low empirical bias in a variety of generative settings.
Jah, A.; Ngesa, O.; Wamwea, C.; Ngunyi, A.
Show abstract
Abstract: Compartmental epidemic models conventionally treat the probability of moving between dis ease states as fixed over time, an assumption that sits uneasily with the reality of a pandemic in which lockdowns, mask mandates, vaccination roll-out, and the arrival of new variants continually reshape transmission. This paper develops a time-inhomogeneous Markov chain framework for the Susceptible-Exposed-Infectious-Removed (SEIR) process, in which each transition probability pab(t) is allowed to vary with calendar time while respecting the struc tural zeros implied by the SEIR compartmental flow. We derive the constrained maximum likelihood estimator of pab(t) under these structural constraints, establish its finite sample efficiency, asymptotic normality, and Wilson score confidence intervals, and construct a like lihood ratio test of the null hypothesis that a compartments exit probability is constant over time. We further propose a stochastic machine learning hybrid extension in which the raw, kernel smoothed transition probabilities are regressed on policy and mobility covariates using both a logistic generalized linear model and a random forest, allowing the framework to attribute time-inhomogeneity to observable interventions. The methodology is applied to a compiled daily, district level COVID-19 surveillance panel for Sierra Leone spanning March 2020 to December 2023 (16 districts, 1,401 days). The likelihood ratio test rejects time-homogeneity of the exposed to infectious transition in 15 of 16 districts and of the infectious-to-removed transition in 8 of 16 districts ( = 0.05), and the covariate augmented logistic model achieves an out of sample Brier score roughly 76 times smaller than a time homogeneous pooled baseline, with healthcare capacity and the time trend emerging as the most influential predictors in the random-forest component. These results provide statisti cal evidence that time-inhomogeneous, covariate informed Markov models offer a materially better description of district-level COVID-19 transmission in Sierra Leone than classical time homogeneous compartmental models, with implications for sub-national outbreak monitoring in resource constrained settings
Bradley, V. C.; Nichols, T. E.
Show abstract
The UK Biobank is a national prospective study of half a million participants between the ages of 40 and 69 at the time of recruitment between 2006 and 2010, established to facilitate research on diseases of aging. The imaging cohort is a subset of UK Biobank participants who have agreed to undergo extensive additional imaging assessments. However, Fry et al. (2017) finds evidence of "healthy volunteer bias" in the UK Biobank - participants are less likely to smoke, be obese, consume alcohol daily than the target population of UK adults. Here we examine selection bias in the UK Biobank imaging cohort. We address two common misconceptions: first, that study size can compensate for bias in data collection, and second that selection bias does not affect estimates of associations, which are the primary interest of the UK Biobank. We introduce inverse probability weighting (IPW) as an approach commonly used in survey research that can be used to address selection bias in volunteer health studies like the UK Biobank. We discuss 6 such methods - five existing and one novel -, assess relative performance in simulation studies, and apply them to the UK Biobank imaging cohort. We find that our novel method, BART for predicting the probability of selection combined with raking, performs well relative to existing methods, and helps alleviate selection bias in the UK Biobank imaging cohort.
Qian, J.; Tanigawa, Y.; Li, R.; Tibshirani, R.; Rivas, M. A.; Hastie, T.
Show abstract
In high-dimensional regression problems, often a relatively small subset of the features are relevant for predicting the outcome, and methods that impose sparsity on the solution are popular. When multiple correlated outcomes are available (multitask), reduced rank regression is an effective way to borrow strength and capture latent structures that underlie the data. Our proposal is motivated by the UK Biobank population-based cohort study, where we are faced with large-scale, ultrahigh-dimensional features, and have access to a large number of outcomes (phenotypes): lifestyle measures, biomarkers, and disease outcomes. We are hence led to fit sparse reduced-rank regression models, using computational strategies that allow us to scale to problems of this size. We use an iterative algorithm that alternates between solving the sparse regression problem and solving the reduced rank decomposition. For the sparse regression component, we propose a scalable iterative algorithm based on adaptive screening that leverages the sparsity assumption and enables us to focus on solving much smaller sub-problems. The full solution is reconstructed and tested via an optimality condition to make sure it is a valid solution for the original problem. We further extend the method to cope with practical issues such as the inclusion of confounding variables and imputation of missing values among the phenotypes. Experiments on both synthetic data and the UK Biobank data demonstrate the effectiveness of the method and the algorithm. We present multiSnpnet package, available at http://github.com/junyangq/multiSnpnet that works on top of PLINK2 files, which we anticipate to be a valuable tool for generating polygenic risk scores from human genetic studies.
Shao, L.; Pohl, K. M.; Thompson, W. K.
Show abstract
The Synthetic Control Method (SCM) and its interactive factor model generalizations (GSC) are powerful for estimating causal effects from panel data but are not easily applied when follow-up is irregular or sparse, common features of biomedical cohorts. We develop a Bayesian functional extension of GSC that treats each units outcome path as a smooth latent trajectory and accommodates unequally spaced measurements. Trajectories are approximated using Functional Principal Components Analysis (FPCA), providing a data-driven basis that captures dominant patterns with minimal shape assumptions while borrowing strength across individuals. Within this representation, we learn unit and time latent factors jointly with FPCA scores from the control data, construct counterfactual trajectories for treated units, and quantify uncertainty via the posterior. Identification relies on a latent-factor/weak-trend condition and overlap of controls and treated units in the functional score space. Simulation studies varying donor pool and treated unit size and sampling density show that the proposed approach (a.k.a GSC-FPCA) yields low bias when sampling is irregular or sparse, with well-calibrated interval coverage across a broad range of scenarios. We apply the method to longitudinal neuroimaging data from the National Consortium on Alcohol and Neurodevelopment in Adolescence - Adulthood (NCANDA-A) study to estimate the effect of adolescent binge drinking on subsequent brain volumes. Leveraging from 1 to 9 observed time points per participant, GSC-FPCA produces stable counterfactuals and detects a negative impact on gray-matter volumes with sustained high levels of binge drinking. Our results demonstrate that embedding GSC within a functional framework enables robust causal inference in biomedical applications characterized by irregularly-spaced visits, limited observations, and complex outcome dynamics.
Yuan, S.; Zou, F.; Zou, B.
Show abstract
Lung transplantation programs must decide when bilateral lung transplantation (BLT) offers meaningful functional benefit over single lung transplantation (SLT). Because donor and recipient characteristics jointly shape outcomes, the BLT-SLT contrast may differ across patients. However, analyzing observational registries poses a statistical challenge: apparent subgroup differences can be artifacts of complex confounding, while true heterogeneity can be missed or poorly quantified. Using a large national registry, we investigate whether the BLT effect varies across recipients and identify clinically relevant profiles of benefit using post-transplant lung function measured by forced expiratory volume in 1 second (FEV1). We develop deepHTL, a framework that tests for treatment effect heterogeneity and estimates how the BLT-SLT effect varies with patient features. In extensive simulations designed to resemble registry-like confounding, deepHTL controls false positives for detecting heterogeneity and yields more accurate individualized effect estimates than common machine learning methods. In the lung transplant cohort, we find strong evidence of heterogeneity in the BLT-SLT effect on FEV1: younger, lower risk recipients with better baseline status show the largest FEV1 gains from BLT, whereas older, higher risk candidates exhibit diminished marginal benefit. These findings provide statistically grounded guidance for patient selection and allocation of scarce donor organs.
Jiang, C.; Fang, F.; Talbot, D.; Schnitzer, M.
Show abstract
The Test-Negative Design (TND), which involves recruiting care-seeking individuals who meet predefined clinical case criteria, offers valid statistical inference for Vaccine Effectiveness (VE) using data collected through passive surveillance, making it cost-efficient and timely. Infectious disease epidemiology often involves interference, where the treatment and/or outcome of one individual can affect the outcomes of others, rendering standard causal estimands ill-defined; ignoring such interference can bias VE evaluation and lead to ineffective vaccination policies. This article addresses the estimation of causal estimands for VE in the presence of partial interference using TND samples. Partial interference means that the vaccination of units within the same group/cluster may influence the outcomes of other members of the cluster. We define the population direct, spillover, total, and overall effects using the geometric risk ratio, which are identifiable under TND sampling. We investigate various stochastic policies for vaccine allocation in a counterfactual scenario, and identify policy-relevant VE causal estimands. We propose inverse-probability weighted (IPW) estimators for estimating the policy-relevant VE causal estimands with partial interference under the TND, and explore the statistical properties of these estimators.
Taylor, A. R.; Foo, Y. S.; White, M. T.
Show abstract
Malaria is a major public health concern. Among the two most important causes of human malaria, Plasmodium vivax is the hardest to eliminate, owing largely to its ability to relapse (cause recurrent blood-stage infection following the activation of dormant liver-stage parasites called hypnozoites). Recurrent vivax malaria can also follow a failure to treat a preceding blood-stage infection (recrudescence) and, in endemic settings, a new infectious mosquito bite (reinfection). Understanding the cause of recurrent vivax malaria is critical for disease control efforts, e.g., to estimate the efficacy of a hypnozoiticidal drug, relapse needs to be separated from reinfection and recrudescence. In this report we describe the Pv3R statistical model designed to estimate using P. vivax genetic data the probability that a recurrent P. vivax blood-stage infection is a relapse, reinfection or recrudescence (the Pv 3Rs).
Nemcova, B.; Goldstein, I. H.; Sebastian, J.; Minin, V. M.; Bracher, J.
Show abstract
Time-varying effective reproductive numbers of infectious diseases are commonly estimated using renewal equation models. In the widely applied R package EpiEstim and various related tools, this approach is combined with a Poisson distributional assumption. This has been criticized on various occasions, mostly on grounds of general model realism or a desire to estimate overdispersion parameters. Here we argue that an important issue arising from the Poisson assumption is that inference about the effective reproductive number becomes overconfident in presence of overdispersion. By how much standard errors are underestimated follows in a straightforward manner from theory on generalized linear models. We therefore recommend to replace the Poisson assumption by quasi-Poisson or negative binomial extensions, and contrast their respective properties. We illustrate our arguments in detailed simulation studies and three examples of case studies of Ebola, pandemic influenza and COVID-19.
Putter, H.; Goeman, J.; Wallinga, J.
Show abstract
Compartmental models based on ordinary differential equations (ODEs) quantifying the interactions between susceptible, infectious, and recovered individuals within a population have played an important role in infectious disease modeling. The aim of the present paper is to explain the link between stochastic epidemic models based on the susceptible-infectious-recovered (SIR) model, and methods from survival analysis. We illustrate how standard software for survival analysis in the statistical language R can be used to estimate pivotal parameters in the stochastic SIR model in the very much idealized situation where the epidemic is completely observed. Extensions incorporating interventions, age structure and heterogeneity are explored and illustrated.
Beaton, D.; Sunderland, K. M.; ADNI, ; Levine, B.; Mandzia, J.; Masellis, M.; Swartz, R. H.; Troyer, A. K.; ONDRI, ; Binns, M. A.; Abdi, H.; Strother, S. C.
Show abstract
The minimum covariance determinant (MCD) algorithm is one of the most common techniques to detect anomalous or outlying observations. The MCD algorithm depends on two features of multivariate data: the determinant of a matrix (i.e., geometric mean of the eigenvalues) and Mahalanobis distances (MD). While the MCD algorithm is commonly used, and has many extensions, the MCD is limited to analyses of quantitative data and more specifically data assumed to be continuous. One reason why the MCD does not extend to other data types such as categorical or ordinal data is because there is not a well-defined MD for data types other than continuous data. To address the lack of MCD-like techniques for categorical or mixed data we present a generalization of the MCD. To do so, we rely on a multivariate technique called correspondence analysis (CA). Through CA we can define MD via singular vectors and also compute the determinant from CAs eigenvalues. Here we define and illustrate a generalized MCD on categorical data and then show how our generalized MCD extends beyond categorical data to accommodate mixed data types (e.g., categorical, ordinal, and continuous). We illustrate this generalized MCD on data from two large scale projects: the Ontario Neurodegenerative Disease Research Initiative (ONDRI) and the Alzheimers Disease Neuroimaging Initiative (ADNI), with genetics (categorical), clinical instruments and surveys (categorical or ordinal), and neuroimaging (continuous) data. We also make R code and toy data available in order to illustrate our generalized MCD.
Fitzgerald, T.; Jones, A.; Engelhardt, B. E.
Show abstract
Single-cell RNA sequencing (scRNA-seq) technologies allow for the study of gene expression in individual cells. Often, it is of interest to understand how transcriptional activity is associated with cell-specific covariates, such as cell type, genotype, or measures of cell health. Traditional approaches for this type of association mapping assume independence between the outcome variables (or genes), and perform a separate regression for each. However, these methods are computationally costly and ignore the substantial correlation structure of gene expression. Furthermore, count-based scRNA-seq data pose challenges for traditional models based on Gaussian assumptions. We aim to resolve these issues by developing a reduced-rank regression model that identifies low-dimensional linear associations between a large number of cell-specific covariates and high-dimensional gene expression readouts. Our probabilistic model uses a Poisson likelihood in order to account for the unique structure of scRNA-seq counts. We demonstrate the performance of our model using simulations, and we apply our model to a scRNA-seq dataset, a spatial gene expression dataset, and a bulk RNA-seq dataset to show its behavior in three distinct analyses. We show that our statistical modeling approach, which is based on reduced-rank regression, captures associations between gene expression and cell- and sample-specific covariates by leveraging low-dimensional representations of transcriptional states.
Camirand-Lemyre, F.; Domingue, M.-P.; Morissette, J.-P.; Burgun, A.; Ethier, J.-F.
Show abstract
Life sciences research increasingly relies on variables held by different entities, such as clinical, laboratory, environmental, and genomic data. Due to legal, ethical, and social acceptability constraints, these data often cannot be shared across organizations holding them. As a result, they cannot be pooled, and analyses must be conducted within the framework of vertically partitioned data. Supporting such analyses requires methods that protect privacy. However, the mere fact that line-level data are not exchanged should not be mistaken for true privacy protection. We introduce VALORIS (Vertically partitioned Analytics under the LOgistic Regression model for Inference in Statistics), a novel method that enables statistical inference under a logistic regression model without disclosing any individual-level data--including the outcome variable. VALORIS is a practical, communication-efficient algorithm that requires no third-party coordinator. Most importantly, it includes a novel framework for evaluating privacy, allowing users to distinguish among different levels of privacy preservation.
Marchi, H.; Schmiegel, S.; Fuchs, C.; Schamberger, T.
Show abstract
The validation of methods is an integral part of statistical research, defining conditions under which methods can yield reliable results. When validation is carried out empirically, it requires a solid data basis that allows the control and management of relevant characteristics. In particular, depending on the context, the data must meet specific requirements regarding, e. g. sample size, dimensionality, completeness and underlying dependency structures. Real-world data often fails to meet these requirements, particularly in medical contexts where availability or the right to publish is additionally restricted through privacy regulations. For this reason, synthetic data is an effective alternative for method validation. Generating synthetic data is particularly demanding if it is required to precisely mirror complex dependence structures while simultaneously controlling certain target characteristics. In our work, we address the medical context of simultaneously co-occurring diseases, where symptoms may overlap or conflict. We seek to generate synthetic data supporting the simulation-based validation of statistical methods which are able to predict holistic disease pictures based on patient information, including the accounting for comorbidities. We introduce a four-step framework in which we (I) generate patient covariates such as symptoms; (II) connect this patient information to predictors for the single or joint occurrence of diseases; (III) transform the predictors into disease probabilities or disease scores; (IV) convert the probabilities or scores into disease occurrence. Within each of these steps, we outline several alternatives that allow different forms of modeling the overall dependence structure. We apply our framework to a case study informed by real-world data: to the context of pain-causing diseases which share certain similarities in their clinical presentations, and which can occur either individually or jointly. In this analysis, we employ five combinations of methodological alternatives within the data generation steps. This allows us to evaluate the data generation approaches with respect to their ability to achieve the defined target characteristics, and to demonstrate strengths and weaknesses as well as specific suitability. The proposed simulation framework is broadly applicable beyond the specific use case and the medical context. The approaches are designed to be accessible and adjustable through varying input settings, enabling users to tailor the data generation to their specific needs. This way, our work provides researchers with a flexible framework for generating synthetic validation data that aligns with the methodological requirements of their studies.
Connor, G.; Pesta, B. J.
Show abstract
Admixture regression methodology exploits the natural experiment of random mating between individuals with different ancestral backgrounds to infer the environmental and genetic components to trait variation across racial and ethnic groups. This paper provides a statistical framework for admixture regression based on the linear polygenic index model and applies it to neuropsychological performance data from the Adolescent Brain Cognitive Development (ABCD) database. We develop and apply a new test of the differential impact of multi-racial identities on trait variation, an orthogonalization procedure for added explanatory variables, and a partially linear semiparametric functional form. We find a statistically significant genetic component to neuropsychological performance differences across racial identities, and find some possible evidence of nonlinearity in the link between admixture and neuropsychological performance scores in the ABCD data.
Yang, Y.; Zhu, X.
Show abstract
Before applying flexible nonparametric models such as a generalized additive model (GAM), it is natural to ask whether a simpler parametric form suffices. To address this question, we develop TAPS (Test for Arbitrary Parametric Structure), a framework that integrates estimation and hypothesis testing to evaluate whether a prespecified parametric form adequately captures a covariate effect in a GAM. TAPS accommodates diverse structures, including linearity, piecewise linearity with changepoints, discontinuities with jumps, among many others. It is implemented in the R package mgcv.taps built directly on mgcv, enabling seamless adoption, broad outcome support, and scalability to biobank-scale data. Using UK Biobank data, we analyze 38 continuous and 8 binary traits to investigate two scientific questions: does the effect of a polygenic risk score (PRS) vary with age beyond a linear interaction, and does retirement at age 65 modify this age-varying effect? We find that age-varying PRS effects are common and often strongly non-linear, and that retirement at 65 significantly modifies these effects for five traits after multiple-testing correction.
Gill, C. C.; Marchini, J.
Show abstract
Disease etiology may be better understood through the study of gene expression in four dimensional (4D) experiments that consist of measurements on multiple individuals, genes, tissues and under multiple conditions or through time. We have developed a sparse Bayesian four dimensional tensor decomposition method aimed at uncovering latent components or gene networks that could be linked to genetic variation. We used a Variational Bayes algorithm to fit the model which provides fast and accurate analysis. In this brief note we illustrate the utility of the method using simulated datasets, and show that when 4D data is available our method shows improved performance in estimating the true structure in the dataset, when compared to using a 3D method on a single slice of the 4D dataset. We also compare the results of the 4D method to that of the 3D method on a suitable unfolding of the dataset, demonstrating that similar performance is observed in this case, while the 4D method accurately recovers the additional structure in the data. We provide software that implements the method in R.
Butzin-Dozier, Z.; Qiu, S.; Hubbard, A. E.; Shi, J.; van der Laan, M.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWUnderstanding treatment effects on health-related outcomes using real-world data requires defining a causal parameter and imposing relevant identification assumptions to translate it into a statistical estimand. Semiparametric methods, like the targeted maximum likelihood estimator (TMLE), have been developed to construct asymptotically linear estimators of these parameters. To further establish the asymptotic efficiency of these estimators, two conditions must be met: 1) the relevant components of the data likelihood must fall within a Donsker class, and 2) the estimates of nuisance parameters must converge to their true values at a rate faster than n-1/4. The Highly Adaptive LASSO (HAL) satisfies these criteria by acting as an empirical risk minimizer within a class of cadlag functions with a bounded sectional variation norm, which is known to be Donsker. HAL achieves the desired rate of convergence, thereby guaranteeing the estimators asymptotic efficiency. The function class over which HAL minimizes its risk is flexible enough to capture realistic functions while maintaining the conditions for establishing efficiency. Additionally, HAL enables robust inference for non-pathwise differentiable parameters, such as the conditional average treatment effect (CATE) and causal dose-response curve, which are important in precision health. While these parameters are often considered in machine learning literature, these applications typically lack proper statistical inference. HAL addresses this gap by providing reliable statistical uncertainty quantification that is essential for informed decision-making in health research.