Back

Mixture Network Regularized Generalized Linear Model with Feature Selection

Li, K.; Wang, X.; Kuan, P. F.

2019-06-21 bioinformatics
10.1101/678029 bioRxiv
Show abstract

High dimensional genomics data in biomedical sciences is an invaluable resource for constructing statistical prediction models. With the increasing knowledge of gene networks and pathways, such information can be utilized in the statistical models to improve prediction accuracy and enhance model interpretability. However, in certain scenarios the network structure may only be partially known or subject to inaccuracy. Thus, the performance of statistical models incorporating such network structure may be compromised. In this paper, we propose a weighted sparse network learning method by optimally combining a data driven network with sparsity property to prior known or partially known network to address this issue. We show that our proposed model attains the oracle property and achieves a parsimonious structure in high dimensional setting for different types of outcomes including continuous, binary and survival data. Simulations studies show that our proposed model is robust and outperforms existing methods. Case study on melanoma gene expression further demonstrates that our proposed model achieves good operating characteristics in identifying informative genes and predicting survival risk. An R package glmaag implementing our method is available on the Comprehensive R Archive Network (CRAN).

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.