Back

Biometrics

Oxford University Press (OUP)

All preprints, ranked by how well they match Biometrics's content profile, based on 23 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
zifalsnm: Zero-Inflated Bayesian factor analysis model with skew-normal priors for modeling microbiome data

Panchasara, S.; Jankowski, H.; McGregor, K.

2025-12-10 genetics 10.64898/2025.12.07.692834 medRxiv
Top 0.1%
45.6%
Show abstract

MotivationAdvancements in next-generation sequencing have transformed our understanding of host-microbe interactions, revealing links between microbial composition and chronic conditions such as obesity, diabetes, IBD, and others. However, the analysis of microbiome data is complex due to its unique statistical characteristics. One primary objective is to achieve effective dimension reduction to manage high dimensionality while simultaneously accounting for the datas compositional nature and zero inflation. Although existing probabilistic models provide frameworks for composition estimation, they are often based on the assumption that log-ratio-transformed compositions are normally distributed. This assumption is problematic, as it often fails to capture the significant skewness inherent in these transformed compositions. ResultsWe propose a new model called the Zero-Inflated Factor Analysis Logistic Skew-Normal Multinomial (ZIFA-LSNM) model : a comprehensive Bayesian hierarchical framework designed to address the statistical challenges of microbiome data. ZIFA-LSNM integrates a zero-inflation component to handle excess zeros, employs factor analysis for dimensionality reduction, and, critically, utilizes skew-normal priors on the latent factors to explicitly model data asymmetry. Posterior inference is performed using a scalable and efficient variational inference algorithm. Through simulation studies and real data analysis, the ZIFA-LSNM model has shown to demonstrate superior performance in parameter recovery and composition estimation compared to its Gaussian-based counterparts. Availability and Implementationzifalsnm is implemented in a freely available R package: https://github.com/SaurabhP-MS/zifalsnm.git Supplementary InformationSupplementary material is available with this article.

2
Identifying Effect Modification of Latent Population Characteristics on Risk Factors with a Sparse Varying Coefficient Regression

Wang, R.; Fang, L.; Wang, Y.; Jin, J.

2024-12-05 genetics 10.1101/2024.11.30.626101 medRxiv
Top 0.1%
45.6%
Show abstract

Leveraging observational data to understand the associations between risk factors and disease outcomes and conduct disease risk prediction is a common task in epidemiology. While traditional linear regression and other machine learning models have been extensively implemented for this task, the associations between risk factors and disease outcomes are typically deemed fixed. In many cases, however, such associations may vary by some underlying features of the individuals, which may involve certain subpopulation characteristics and environmental factors. While data for these latent features may not be available, the observed data on risk factors may have captured some proportion of the variation in these features. Thus extracting latent factors from risk factors and incorporating this effect modification into the model may better capture the underlying data structure and improve inference. We develop a novel regression model with some coefficients varying as functions of latent features extracted from the risk factors. We have demonstrated the superiority of our approach in various data settings via simulation studies. An application on a dataset for lung cancer patients from The Cancer Genome Atlas (TCGA) Program showed that our approach led to a 6% - 118% increase in (AUC-0.5) for distinguishing between different lung cancer stages compared to the classic lasso and elastic net regressions and identified interesting latent effect modifications associated with certain gene pathways.

3
Transfer Learning for Survival-based Clustering of Predictors with an Application to TP53 Mutation Annotation

Liu, X.; Yan, H.; Shi, H.; Montellier, E.; Chi, E. C.; Hainaut, P.; Wang, W.

2025-10-06 genetics 10.1101/2025.10.06.680732 medRxiv
Top 0.1%
45.2%
Show abstract

TP53 is the most frequently mutated gene in human cancers, and germline mutations in TP53 cause Li-Fraumeni syndrome (LFS), a hereditary predisposition to diverse cancers. Accurate annotation of TP53 mutations based on their survival effects is critical for informed LFS patient management. Motivated by this need, we develop a new approach for Survival-based Clustering of Predictors (SCP) by identifying homogeneous coefficients in Cox regression. We formulate this task as a fusionpenalized Cox regression problem and provide an efficient computational algorithm. A nonconvex distance-to-set penalty is adopted to facilitate parameter tuning and improve estimation accuracy. To overcome data limitations, we further develop TLSCP, a transfer learning extension that borrows coefficient ranking information from a source dataset under the assumption of similar ranking patterns between source and target. TL-SCP integrates ranking information through weighted rank averaging, allowing flexibility in accommodating cohort heterogeneity while maintaining model simplicity. Simulation studies demonstrate TL-SCPs superior performance over SCP in clustering recovery and coefficient estimation. In the application of TP53 mutation annotation where we utilize non-LFS germline TP53 mutation carriers as a source cohort for the target LFS cohort, TL-SCP identifies biologically meaningful TP53 mutation clusters and offers improved clinical interpretability compared to experiment-based annotations.

4
Robust Inference of Individualized Treatment Effect in Mendelian Randomization

Liang, M.; Wu, R.; Xiao, F.; Li, X.

2026-05-12 genetics 10.64898/2026.05.08.723855 medRxiv
Top 0.1%
30.6%
Show abstract

Mendelian randomization (MR) is widely used to draw causal conclusions in the presence of unmeasured confounding, but most MR analyses focus on average treatment effects and rely on strong assumptions. For precision medicine, the primary target is instead the individualized treatment effect (ITE); yet in MR, such effects are not point-identified under core IV assumptions, and valid inference is particularly challenging. We therefore propose a robust partial identification inference framework for ITE under MR allowing multiple instruments. Under minimal causal assumptions, we derive a sharp inference procedure for the intersection bounds of ITE by adopting a multiplier bootstrap procedure with data-adaptive bootstrap distribution shifting and heterogeneous variance adjustment. In theory, we prove that the proposed method achieves nominal coverage and asymptotic sharpness. Further, we extend the procedure to tolerate possible invalid IVs under a minimal proportion rule assumption by aggregating over instrument subsets while preserving coverage. Simulation studies demonstrate that the proposed methods attain nominal coverage and substantially shorter intervals than existing procedures. We illustrate the framework using data from the Alzheimers Disease Neuroimaging Initiative to assess heterogeneous causal effects of TREM2 expression on Alzheimers disease risk across education-defined subgroups.

5
Estimating Direct and Spillover Vaccine Effectiveness with Partial Interference under Test-Negative Design Sampling

Jiang, C.; Fang, F.; Talbot, D.; Schnitzer, M.

2025-02-25 infectious diseases 10.1101/2025.02.24.25322826 medRxiv
Top 0.1%
30.1%
Show abstract

The Test-Negative Design (TND), which involves recruiting care-seeking individuals who meet predefined clinical case criteria, offers valid statistical inference for Vaccine Effectiveness (VE) using data collected through passive surveillance, making it cost-efficient and timely. Infectious disease epidemiology often involves interference, where the treatment and/or outcome of one individual can affect the outcomes of others, rendering standard causal estimands ill-defined; ignoring such interference can bias VE evaluation and lead to ineffective vaccination policies. This article addresses the estimation of causal estimands for VE in the presence of partial interference using TND samples. Partial interference means that the vaccination of units within the same group/cluster may influence the outcomes of other members of the cluster. We define the population direct, spillover, total, and overall effects using the geometric risk ratio, which are identifiable under TND sampling. We investigate various stochastic policies for vaccine allocation in a counterfactual scenario, and identify policy-relevant VE causal estimands. We propose inverse-probability weighted (IPW) estimators for estimating the policy-relevant VE causal estimands with partial interference under the TND, and explore the statistical properties of these estimators.

6
De-biased sparse canonical correlation for identifying cancer-related trans-regulated genes

Huey, N.; Dutta, D.; Laha, N.

2024-08-19 genetics 10.1101/2024.08.15.608166 medRxiv
Top 0.1%
27.4%
Show abstract

SO_SCPLOWUMMARYC_SCPLOWIn cancer multi-omic studies, identifying the effects of somatic copy number aberrations (CNA) on physically distal gene expressions (trans-associations) can potentially uncover genes critical for cancer pathogenesis. Sparse canonical correlation analysis (SCCA) has emerged as a promising method for identifying associations in high-dimensional settings, owing to its ability to aggregate weaker associations and its improved interpretability. Traditional SCCA lacks hypothesis testing capabilities, which are critical for controlling false discoveries. This limitation has recently been addressed through a bias correction technique that enables calibrated hypothesis testing. In this article, we leverage the theoretical advancements in de-biased SCCA to present a computationally efficient pipeline for multi-omics analysis. This pipeline identifies and tests associations between multi-omics data modalities in biomedical settings, such as the trans-effects of CNA on gene expression. We propose a detailed algorithm to choose the tuning parameters of de-biased SCCA. Applying this pipeline to data on estrogen receptor (ER)-associated CNAs and 10,756 gene expressions from 1,904 breast cancer patients in the METABRIC study, we identified 456 CNAs trans-associated with 256 genes. Among these, 5 genes were identified only through de-biased SCCA and not by the standard pairwise regression approach. Downstream analysis with the 256 genes revealed that these genes were overrepresented in pathways relevant to breast cancer.

7
BMDD: A Probabilistic Framework for Accurate Imputation of Zero-inflated Microbiome Sequencing Data

Zhou, H.; Chen, J.; Zhang, X.

2025-05-12 genomics 10.1101/2025.05.08.652808 medRxiv
Top 0.1%
26.8%
Show abstract

Microbiome sequencing data are inherently sparse and compositional, with excessive zeros arising from biological absence or insufficient sampling. These zeros pose significant challenges for downstream analyses, particularly those that require log-transformation. We introduce BMDD (BiModal Dirichlet Distribution), a novel probabilistic modeling framework for accurate imputation of microbiome sequencing data. Unlike existing imputation approaches that assume unimodal abundance, BMDD captures the bimodal abundance distribution of the taxa via a mixture of Dirichlet priors. It uses variational inference and a scalable expectation-maximization algorithm for efficient imputation. Through simulations and real microbiome datasets, we demonstrate that BMDD outperforms competing methods in reconstructing true abundances and improves the performance of differential abundance analysis. Through multiple posterior samples, BMDD enables robust inference by accounting for uncertainty in zero imputation. Our method offers a principled and computationally efficient solution for analyzing high-dimensional, zero-inflated microbiome sequencing data and is broadly applicable in microbial biomarker discovery and host-microbiome interaction studies. BMDD is available at: https://github.com/zhouhj1994/BMDD. Author SummaryUnderstanding the microbes living in and on our bodies--the microbiome--relies on analyzing complex sequencing data. However, these data often contain many zeros, either because a microbe is truly absent or simply missed due to insufficient sampling. These missing values make it hard to accurately analyze microbial patterns and identify important differences between groups, especially for methods that work on a log scale. To address this, we developed a new method called BMDD that uses a more realistic model to impute the zeros. Unlike existing tools that assume each microbe follows an unimodal abundance distribution, BMDD allows for microbes to follow a bimodal distribution, so they could behave differently in different conditions. It provides not just a single guess, but a range of possible values to better reflect the uncertainty. Our testing shows that BMDD more accurately recovers the true microbial profiles and improves the ability to detect meaningful differences between groups. This method can help researchers gain clearer insights into how the microbiome affects health and disease.

8
Personalized Risk Prediction for Cancer Survivors: A Bayesian Semi-parametric Recurrent Event Model with Competing Outcomes

Nguyen, N. H.; Shin, S. J.; Dodd-Eaton, E. B.; Ning, J.; Wang, W.

2023-03-01 genetics 10.1101/2023.02.28.530537 medRxiv
Top 0.1%
22.9%
Show abstract

Multiple primary cancers are increasingly more frequent due to improved survival of cancer patients. Characteristics of the first primary cancer largely impact the risk of developing subsequent primary cancers. Hence, model-based risk characterization of cancer survivors that captures patient-specific variables is needed for healthcare policy making. We propose a Bayesian semi-parametric framework, where the occurrence processes of the competing cancer types follow independent non-homogeneous Poisson processes and adjust for covariates including the type and age at diagnosis of the first primary. Applying this framework to a historically collected cohort with families presenting a highly enriched history of multiple primary tumors and diverse cancer types, we have derived a suite of age-to-onset penetrance curves for cancer survivors. This includes penetrance estimates for second primary lung cancer, potentially impactful to ongoing cancer screening decisions. Using Receiver Operating Characteristic (ROC) curves, we have validated the good predictive performance of our models in predicting second primary lung cancer, sarcoma, breast cancer, and all other cancers combined, with areas under the curves (AUCs) at 0.89, 0.91, 0.76 and 0.68, respectively. In conclusion, our framework provides covariate-adjusted quantitative risk assessment for cancer survivors, hence moving a step closer to personalized health management for this unique population.

9
A Bayesian Informative Shrinkage Approach for Large-scale Multiple Hypothesis Testing (BISHOT): with Applications in Differential Analysis of Omics Data

Su, Y.; Clark, M. E. J. Z.; Wang, C.

2025-09-16 genetics 10.1101/2025.09.11.675690 medRxiv
Top 0.1%
22.5%
Show abstract

A major goal of many omics studies is to identify differential features, e.g. differentially expressed genes, between experimental groups. When performing differential analysis for a given dataset, relevant information from another platform or species is often available. Incorporating such prior information can help identify features that show consistent differential patterns across platforms or species, which are more likely to reflect shared biological processes, and thereby enhance the robustness and generalizability of the findings. However, existing differential analysis methods typically analyze only the data from the current study and do not leverage prior knowledge about the magnitude or direction of changes from other platforms or species. We address this challenge, and the associated multiple testing problem, using a Bayesian framework that enables the incorporation of prior knowledge obtained from different platforms or species. We propose a new test statistic, Bayesian Credible Ratio (BCR), based on a heteroscedastic global local shrinkage prior, and a new multiple testing criterion, sign-adjusted FDR (SFDR), that emphasize information regarding the direction of the differentially features. We prove that BCR achieves the largest count of sign-based true positives among all legitimate SFDR-controlling methods. Simulation results offer numerical evidence of its advantage compared to an empirical Bayesian method. The approach is demonstrated through the analysis of RNAseq and single-cell RNAseq datasets.

10
Estimation of total mediation effect for a binary trait in a case-control study for high-dimensional omics mediators

Kang, Z.; Chen, L.; Wei, P.; Xu, Z.; Li, C.; Yang, T.

2025-02-02 genomics 10.1101/2025.01.28.635396 medRxiv
Top 0.1%
22.4%
Show abstract

Mediation analysis helps uncover how exposures impact outcomes through intermediate variables. Traditional mean-based total mediation effect measures may suffer from the cancellation of opposite component-wise effects, and existing methods often lack the power to capture weak effects in high-dimensional mediators. Additionally, most high-dimensional mediation analysis methods have focused on continuous outcomes, with limited attention to binary outcomes, particularly in case-control studies. To fill this gap, we propose an R2 total mediation effect measure within the liability framework that offers a clear and intuitive causal interpretation, provides additional insights beyond the mean-based measures, and is invariant to disease prevalence. We develop a cross-fitted, modified Haseman-Elston regression-based estimation procedure tailored for mediation analysis in case-control studies, which can also be applied to cohort studies. Our estimator remains consistent in the presence of non-mediators and weak effects, as demonstrated in extensive simulations. Theoretical justification for consistency is provided under mild conditions and without requiring exact mediator selection. In a case-control substudy of the Womens Health Initiative involving 2150 individuals, we found that many metabolites were mediators with weak effects in the path from BMI to coronary heart disease, and we estimated that 89% (95% CI: 57%-100%) of the BMI-explained variation in underlying CHD liability is mediated by the measured metabolomics. The proposed estimation procedure is implemented in the R package "r2MedCausal", available on GitHub.

11
Estimation of Mediation Effect for High-dimensional Omics Mediators with Application to the Framingham Heart Study

Yang, T.; Niu, J.; Chen, H.; Wei, P.

2019-09-19 genomics 10.1101/774877 medRxiv
Top 0.1%
22.3%
Show abstract

Environmental exposures can regulate intermediate molecular phenotypes, such as gene expression, by different mechanisms and thereby lead to various health outcomes. It is of significant scientific interest to unravel the role of potentially high-dimensional intermediate phenotypes in the relationship between environmental exposure and traits. Mediation analysis is an important tool for investigating such relationships. However, it has mainly focused on low-dimensional settings, and there is a lack of a good measure of the total mediation effect. Here, we extend an R-squared (Rsq) effect size measure, originally proposed in the single-mediator setting, to the moderate- and high-dimensional mediator settings in the mixed model framework. Based on extensive simulations, we compare our measure and estimation procedure with several frequently used mediation measures, including product, proportion, and ratio measures. Our Rsq measure has small bias and variance under the correctly specified model. To mitigate potential bias induced by non-mediators, we examine two variable selection procedures, i.e., iterative sure independence screening and false discovery rate control, to exclude the non-mediators. We evaluate the consistency of the proposed estimation procedures and introduce a resampling-based confidence interval. By applying the proposed estimation procedure, we find that more than half of the aging-related variations in systolic blood pressure can be explained by gene expression profiles in the Framingham Heart Study.

12
BayICE: A hierarchical Bayesian deconvolution model with stochastic search variable selection

Tai, A.-S.; Tseng, G.; Hsieh, W.-P.

2019-08-12 genomics 10.1101/732743 medRxiv
Top 0.1%
19.2%
Show abstract

Gene expression deconvolution is a powerful tool for exploring the microenvironment of complex tissues comprised of multiple cell groups using transcriptomic data. Characterizing cell activities for a particular condition has been regarded as a primary mission against diseases. For example, cancer immunology aims to clarify the role of the immune system in the progression and development of cancer through analyzing the immune cell components of tumors. To that end, many deconvolution methods have been proposed for inferring cell subpopulations within tissues. Nevertheless, two problems limit the practicality of current approaches. First, all approaches use external purified data to preselect cell type-specific genes that contribute to deconvolution. However, some types of cells cannot be found in purified profiles and the genes specifically over- or under-expressed in them cannot be identified. This is particularly a problem in cancer studies. Hence, a preselection strategy that is independent from deconvolution is inappropriate. The second problem is that existing approaches do not recover the expression profiles of unknown cells present in bulk tissues, which results in biased estimation of unknown cell proportions. Furthermore, it causes the shift-invariant property of deconvolution to fail, which then affects the estimation performance. To address these two problems, we propose a novel deconvolution approach, BayICE, which employs hierarchical Bayesian modeling with stochastic search variable selection. We develop a comprehensive Markov chain Monte Carlo procedure through Gibbs sampling to estimate cell proportions, gene expression profiles, and signature genes. Simulation and validation studies illustrate that BayICE outperforms existing deconvolution approaches in estimating cell proportions. Subsequently, we demonstrate an application of BayICE in the RNA sequencing of patients with non-small cell lung cancer. The model is implemented in the R package \"BayICE\" and the algorithm is available for download.

13
Methods for detection of clusters of observations with an outlying correlation coefficient value

Desmet, L.; Venet, D.; Trotta, L.; Burzykowski, T.; Buyse, M.

2020-10-14 health systems and quality improvement 10.1101/2020.10.12.20211128 medRxiv
Top 0.1%
19.1%
Show abstract

Multivariate datasets with a clustered structure are the natural framework for, e.g., multicentre clinical trials. We propose a number of methods aimed at detecting clusters with outlying correlation coefficients. While the methods can be used in a variety of settings, we focus mainly on their application to central statistical monitoring of clinical trials. In particular, we consider the issue of detecting centers (or other clusters of patients such as regions) with outlying correlation coefficients for bivariate data in a multicenter clinical trial. It appears that, in that context, the proposed methods perform well, as we show by using a simulation study and a number of real life datasets.

14
On Multiply Robust Mendelian Randomization (MR2) With Many Invalid Genetic Instruments

Sun, B.; Liu, Z.; Tchetgen Tchetgen, E.

2021-10-26 epidemiology 10.1101/2021.10.21.21265317 medRxiv
Top 0.1%
18.9%
Show abstract

Mendelian randomization (MR) is a popular instrumental variable (IV) approach, in which genetic markers are used as IVs. In order to improve efficiency, multiple markers are routinely used in MR analyses, leading to concerns about bias due to possible violation of IV exclusion restriction of no direct effect of any IV on the out-come other than through the exposure in view. To address this concern, we introduce a new class of Multiply Robust MR (MR2) estimators that are guaranteed to remain consistent for the causal effect of interest provided that at least one genetic marker is a valid IV without necessarily knowing which IVs are invalid. We show that the proposed MR2 estimators are a special case of a more general class of estimators that remain consistent provided that a set of at least k{dagger} out of K candidate instrumental variables are valid, for k{dagger}[≤] K set by the analyst ex ante, without necessarily knowing which IVs are invalid. We provide formal semiparametric theory supporting our results, and characterize the semiparametric efficiency bound for the exposure causal effect which cannot be improved upon by any regular estimator with our favorable robustness property. We conduct extensive simulation studies and apply our methods to a large-scale analysis of UK Biobank data, demonstrating the superior empirical performance of MR2 compared to competing MR methods.

15
Highly adaptive LASSO: Machine learning that provides valid nonparametric inference in realistic models

Butzin-Dozier, Z.; Qiu, S.; Hubbard, A. E.; Shi, J.; van der Laan, M.

2024-10-19 epidemiology 10.1101/2024.10.18.24315778 medRxiv
Top 0.1%
18.2%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWUnderstanding treatment effects on health-related outcomes using real-world data requires defining a causal parameter and imposing relevant identification assumptions to translate it into a statistical estimand. Semiparametric methods, like the targeted maximum likelihood estimator (TMLE), have been developed to construct asymptotically linear estimators of these parameters. To further establish the asymptotic efficiency of these estimators, two conditions must be met: 1) the relevant components of the data likelihood must fall within a Donsker class, and 2) the estimates of nuisance parameters must converge to their true values at a rate faster than n-1/4. The Highly Adaptive LASSO (HAL) satisfies these criteria by acting as an empirical risk minimizer within a class of cadlag functions with a bounded sectional variation norm, which is known to be Donsker. HAL achieves the desired rate of convergence, thereby guaranteeing the estimators asymptotic efficiency. The function class over which HAL minimizes its risk is flexible enough to capture realistic functions while maintaining the conditions for establishing efficiency. Additionally, HAL enables robust inference for non-pathwise differentiable parameters, such as the conditional average treatment effect (CATE) and causal dose-response curve, which are important in precision health. While these parameters are often considered in machine learning literature, these applications typically lack proper statistical inference. HAL addresses this gap by providing reliable statistical uncertainty quantification that is essential for informed decision-making in health research.

16
On a Unifying ‘Reverse’ Regression for Robust Association Studies and Allele Frequency Estimation with Related Individuals

Zhang, L.; Sun, L.

2019-06-04 genetics 10.1101/470328 medRxiv
Top 0.1%
18.1%
Show abstract

For genetic association studies with related individuals, standard linear mixed-effect model is the most popular approach. The model treats a complex trait (phenotype) as the response variable while a genetic variant (genotype) as a covariate. An alternative approach is to reverse the roles of phenotype and genotype. This class of tests includes quasi-likelihood based score tests. In this work, after reviewing these existing methods, we propose a general, unifying reverse regression framework. We then show that the proposed method can also explicitly adjust for potential departure from Hardy-Weinberg equilibrium. Lastly, we demonstrate the additional flexibility of the proposed model on allele frequency estimation, as well as its connection with earlier work of best linear unbiased allele-frequency estimator. We conclude the paper with supporting evidence from simulation and application studies.

17
Spatial IMIX: A Mixture Model Approach to Spatially Correlated Multi-Omics Data Integration

Wang, Z.; Czerniak, B.; Wei, P.

2023-07-17 genomics 10.1101/2023.07.15.549148 medRxiv
Top 0.1%
18.0%
Show abstract

Spatial high-throughput omics data allow scientists to study gene activity in a tissue sample and map where it occurs at the same time. This enables the possibility to investigate important early cancer-initiating events occur in normal-appearing tissue and gene activities that progress and carry through tumor tissue, as defined by "field effect." The "field effect" genes are differentially expressed or methylated genes in the spatially resolved high-dimensional datasets with respect to the pathology subtype in each geographical sample across the tissue region. Current statistical methods for spatially resolved genomics data focus on the association of omics data with spatial coordinates without being able to incorporate and test for the association with the sample subtypes. In addition, analytical methods are underdeveloped for spatially resolved multi-omics data integration. We propose a novel statistical frame-work spatial IMIX to integratively analyze spatially resolved high-dimensional multi-omics data associated with a specific trait, such as sample subtypes while modeling the spatial correlations between samples and the inter-data-type correlations between omics data simultaneously. Through extensive simulations, spatial IMIX demonstrated well-controlled type I error, great power by relaxing the independence assumptions between data types, model selection features, and the ability to control FDR across data types. Data applications to a geographically annotated tissue area of bladder cancer discovered cancer-initiating gene activities and revealed interesting fundamental biological mechanisms through path-way analysis. We have implemented our method in R package spatialimix available at https://github.com/ziqiaow/spatialimix.

18
A Zero-Inflated Hierarchical Generalized Transformation Model to Address Non-Normality in Spatially-Informed Cell-Type Deconvolution

Melton, H. J.; Bradley, J. R.; Wu, C.

2026-03-06 genomics 10.1101/2024.06.24.600480 medRxiv
Top 0.1%
17.9%
Show abstract

Oral squamous cell carcinomas (OSCC), the predominant head and neck cancer, pose significant challenges due to late-stage diagnoses and low five-year survival rates. Spatial transcriptomics offers a promising avenue to decipher the genetic intricacies of OSCC tumor microenvironments. In spatial transcriptomics, Cell-type deconvolution is a crucial inferential goal; however, current methods fail to consider the high zero-inflation present in OSCC data. To address this, we develop a novel zero-inflated version of the hierarchical generalized transformation model (ZI-HGT) and apply it to the Conditional AutoRegressive Deconvolution (CARD) for cell-type deconvolution. The ZI-HGT serves as an auxiliary Bayesian technique for CARD, reconciling the highly zero-inflated OSCC spatial transcriptomics data with CARDs normality assumption. The combined ZI-HGT + CARD framework achieves enhanced cell-type deconvolution accuracy and quantifies uncertainty in the estimated cell-type proportions. We demonstrate the superior performance through simulations and analysis of the OSCC data. Furthermore, our approach enables the determination of the locations of the diverse fibroblast population in the tumor microenvironment, critical for understanding tumor growth and immunosuppression in OSCC.

19
High-dimensional causal mediation analysis by partial sum statistic and sample splitting strategy in imaging genetics application

Chang, H.-C.; Fang, Y.; Gorczyca, M. T.; Batmanghelich, K.; Tseng, G. C.

2024-06-24 radiology and imaging 10.1101/2024.06.23.24309362 medRxiv
Top 0.1%
15.0%
Show abstract

Causal mediation analysis provides a systematic approach to explore the causal role of one or more mediators in the association between exposure and outcome. In omics or imaging data analysis, mediators are often high-dimensional, which brings new statistical challenges. Existing methods either violate causal assumptions or fail in interpretable variable selection. Additionally, mediators are often highly correlated, presenting difficulties in selecting and prioritizing top mediators. To address these issues, we develop a framework using Partial Sum Statistic and Sample Splitting Strategy, namely PS5, for high-dimensional causal mediation analysis. The method provides a powerful global mediation test satisfying causal assumptions, followed by an algorithm to select and prioritize active mediators with quantification of individual mediation contributions. We demonstrate its accurate type I error control, superior statistical power, reduced bias in mediation effect estimation, and accurate mediator selection using extensive simulations of varying levels of effect size, signal sparsity, and mediator correlations. Finally, we apply PS5 to an imaging genetics dataset of chronic obstructive pulmonary disease (COPD) patients (N=8,897) in the COPDGene study to examine the causal mediation role of lung images (p=5,810) in the associations between polygenic risk score and lung function and between smoking exposure and lung function, respectively. Both causal mediation analyses successfully estimate the global indirect effect and detect mediating image regions. Collectively, we find a region in the lower lobe of the right lung with a strong and concordant mediation effect for both genetic and environmental exposures. This suggests that targeted treatment toward this region might mitigate the severity of COPD due to genetic and smoking effects.

20
Faecal shedding models for SARS-CoV-2 RNA amongst hospitalised patients and implications for wastewater-based epidemiology

Hoffmann, T.; Alsing, J.

2021-03-17 infectious diseases 10.1101/2021.03.16.21253603 medRxiv
Top 0.1%
12.8%
Show abstract

SummaryThe concentration of SARS-CoV-2 RNA in faeces is not well established, posing challenges for wastewater-based surveillance of COVID-19 and risk assessments of environmental transmission. We develop versatile hierarchical models for faecal RNA shedding and apply them to data collected in six studies. We find that the mean number of gene copies per mL of faeces is 1.9 x 106 (2.3 x 105-2.0 x 108 95% credible interval) among unvaccinated hospitalised patients. Using Bayesian model comparison, we find no evidence for a subpopulation of patients who do not shed RNA: limits of quantification can account for negative stool samples. Our models indicate that hospitalised patients represent the tail of the shedding profile with a half-life of 34 hours (28-43 95% credible interval), suggesting that wastewater-based surveillance signals are more indicative of incidence than prevalence and can be a leading indicator of clinical presentation. Shedding among inpatients cannot explain high RNA concentrations observed in wastewater, consistent with more abundant shedding during the early infection course. We show that the models generalise and can predict summary statistics of held-out clinical datasets. However, shedding prior to hospitalisation cannot be constrained due to lack of samples, and information on viral variants was not available.