Back

High-dimensional Bayesian phenotype classification and model selection using genomic predictors

Linder, D. F.; Panchal, V.

2019-09-23 bioinformatics
10.1101/778472 bioRxiv
Show abstract

MotivationIn this paper we describe a Bayesian hierarchical model termed PMMLogit for classification and model selection in high-dimensional settings with binary phenotypes as outcomes. Posterior computation in the logistic model is known to be computationally demanding due to its non-conjugacy with common priors. We combine a Polya-Gamma based data augmentation strategy and use recent results on Markov chain Monte-Carlo (MCMC) techniques to develop an efficient and exact sampling strategy for the posterior computation. We use the resulting MCMC chain for model selection and choose the best combination(s) of genomic variables via posterior model probabilities. Further, a Bayesian model averaging (BMA) approach using the posterior mean, which averages across visited models, is shown to give superior prediction of phenotypes given genomic measurements.\n\nResultsUsing simulation studies, we compared the performance of the proposed method with other popular methods. Simulation results show that the proposed method is quite effective in selecting the true model and has better estimation and prediction accuracy than other methods. These observations are consistent with theoretical results that have been developed in the statistics literature on optimality for this class of priors. Application to two well-known datasets on colon cancer and leukemia identified genes that have been previously reported in the clinical literature to be related to the disease outcomes.\n\nAvailabilitySource code is publicly available on GitHub at https://github.com/v-panchal/PMML.\n\nContactdlinder@augusta.edu\n\nSupplementary informationSupplementary data are available online.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 2%
11.8%
2
Biostatistics
24 papers in training set
Top 0.1%
10.9%
3
The Annals of Applied Statistics
19 papers in training set
Top 0.1%
9.6%
4
BMC Bioinformatics
457 papers in training set
Top 0.8%
8.8%
5
Biometrics
23 papers in training set
Top 0.1%
4.3%
6
PLOS Computational Biology
1863 papers in training set
Top 9%
4.0%
7
BioData Mining
22 papers in training set
Top 0.1%
4.0%
50% of probability mass above
8
Frontiers in Genetics
230 papers in training set
Top 1.0%
3.5%
9
Journal of Computational Biology
48 papers in training set
Top 0.3%
3.2%
10
Statistics in Medicine
40 papers in training set
Top 0.2%
3.2%
11
PLOS ONE
5266 papers in training set
Top 38%
3.2%
12
IEEE/ACM Transactions on Computational Biology and Bioinformatics
38 papers in training set
Top 0.3%
3.2%
13
Frontiers in Artificial Intelligence
20 papers in training set
Top 0.2%
2.4%
14
BMC Genomics
406 papers in training set
Top 3%
2.4%
15
Briefings in Bioinformatics
354 papers in training set
Top 4%
2.1%
16
PeerJ
308 papers in training set
Top 5%
1.9%
17
Journal of Theoretical Biology
162 papers in training set
Top 1%
1.7%
18
Bioinformatics Advances
203 papers in training set
Top 3%
1.3%
19
Scientific Reports
3612 papers in training set
Top 70%
1.0%
20
Computational and Structural Biotechnology Journal
242 papers in training set
Top 7%
0.8%
21
PLOS Genetics
862 papers in training set
Top 12%
0.8%
22
Statistical Methods in Medical Research
11 papers in training set
Top 0.2%
0.6%
23
G3 Genes|Genomes|Genetics
351 papers in training set
Top 4%
0.6%
24
Computer Methods and Programs in Biomedicine
28 papers in training set
Top 1%
0.6%
25
Computers in Biology and Medicine
128 papers in training set
Top 5%
0.6%
26
Frontiers in Systems Biology
10 papers in training set
Top 0.2%
0.6%