Bayesian Shrinkage Priors in Zero-Inflated and Negative Binomial Regression models with Real World Data Applications of COVID-19 Vaccine, and RNA-Seq
Bhattacharyya, A.; mitra, r.; Rai, S.; Pal, S.
Show abstract
BackgroundCount data regression modeling has received much attention in several science fields in which the Poisson, Negative binomial, and Zero-Inflated models are some of the primary regression techniques. Negative binomial regression is applied to modeling count variables, usually when they are over-dispersed. A Poisson distribution is also utilized for counting data where the mean is equal to the variance. This situation is often unrealistic since the distribution of counts will usually have a variance that is not equal to its mean. Modeling it as Poisson distributed leads to ignoring under- or overdispersion, depending on if the variance is smaller or larger than the mean. Also, situations with outcomes having a larger number of zeros such as RNASeq data require Zero-inflated models. Variable selection through shrinkage priors has been a popular method to address the curse of dimensionality and achieve the identification of significant variables. MethodsWe present a unified Bayesian hierarchical framework that implements and compares shrinkage priors in negative-binomial and zero-inflated negative binomial regression models. The key feature is the representation of the likelihood by a Polya-Gamma data augmentation, which admits a natural integration with a family of shrinkage priors. We specifically focus on the Horseshoe, Dirichlet Laplace, and Double Pareto priors. Extensive simulation studies address the efficiency of the model and mean square errors are reported. Further, the models are applied to data sets such as the Covid-19 vaccine, and Covid-19 RNA-Seq data among others. ResultsThe models are robust enough to address variable selection, and MSE decreases as the sample size increases, having lower errors in p > n cases. The noteworthy results showed that the adverse events of Covid-19 vaccines were dependent on age, recovery, medical history, and prior vaccination with a remarkable reduction in MSE of the fitted values. No. of publications of Ph.D. students were dependent on the no. of children, and the no. of articles in the last three years. ConclusionsThe models are robust enough to conduct both variable selections and produce effective fit because of their high shrinkage property and applicability to a broad range of biometric and public health high dimensional problems.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Using a supervised principal components analysis for variable selection in high-dimensional datasets reduces false discovery rates 96%
- Dirichlet distribution parameter estimation with application in microbiome analyses 95%
- A Stability-Enhanced Lasso Approach for Covariate Selection in Non-Linear Mixed Effect Model 95%
Similar papers in this journal
- Genomic prediction using machine learning: A comparison of the performance of regularized regression, ensemble, instance-based and deep learning methods on synthetic and empirical data 95%
- MOSCATO: A Supervised Approach for Analyzing Multi-Omic Single-Cell Data 94%
- Hierarchical non-negative matrix factorization using clinical information for microbial communities. 93%
Similar papers in this journal
- The Burr distribution as a model for the delay between key events in an individual’s infection history 96%
- Model guided trait-specific co-expression network estimation as a new perspective for identifying molecular interactions and pathways 95%
- Covering Hierarchical Dirichlet Mixture Models on binary data to enhance genomic stratifications in Onco-Hematology 95%
Similar papers in this journal
- Two-Stage Multivariate Mendelian Randomization on Multiple Outcomes with Mixed Distributions 94%
- A bivariate zero-inflated negative binomial model and its applications to biomedical settings 94%
- Tight Fit of the SIR Dynamic Epidemic Model to Daily Cases of COVID-19 Reported During the 2021-2022 Omicron Surge in New York City: A Novel Approach 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.