Mixture Network Regularized Generalized Linear Model with Feature Selection
Li, K.; Wang, X.; Kuan, P. F.
Show abstract
High dimensional genomics data in biomedical sciences is an invaluable resource for constructing statistical prediction models. With the increasing knowledge of gene networks and pathways, such information can be utilized in the statistical models to improve prediction accuracy and enhance model interpretability. However, in certain scenarios the network structure may only be partially known or subject to inaccuracy. Thus, the performance of statistical models incorporating such network structure may be compromised. In this paper, we propose a weighted sparse network learning method by optimally combining a data driven network with sparsity property to prior known or partially known network to address this issue. We show that our proposed model attains the oracle property and achieves a parsimonious structure in high dimensional setting for different types of outcomes including continuous, binary and survival data. Simulations studies show that our proposed model is robust and outperforms existing methods. Case study on melanoma gene expression further demonstrates that our proposed model achieves good operating characteristics in identifying informative genes and predicting survival risk. An R package glmaag implementing our method is available on the Comprehensive R Archive Network (CRAN).
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- RCFGL: Rapid Condition adaptive Fused Graphical Lasso and application to modeling brain region co-expression networks 97%
- Extended Graphical Lasso for Multiple Interaction Networks for High Dimensional Omics Data 97%
- Model guided trait-specific co-expression network estimation as a new perspective for identifying molecular interactions and pathways 97%
Similar papers in this journal
- Transfer Learning Models for Bacterial Strain Dissemination Biomarkers using Weighted Non-Parallel Proximal Support Vector Machines 96%
- TISSLET Tissues-based Learning Estimation for Transcriptomics 96%
- DAGBagM: Learning directed acyclic graphs of mixed variables with an application to identify prognostic protein biomarkers in ovarian cancer 95%
Similar papers in this journal
- PALLAS: Penalized mAximum LikeLihood and pArticle Swarms for inference of gene regulatory networks from time series data 96%
- PAN: Personalized Annotation-based Networks for the Prediction of Breast Cancer Relapse 95%
- Supervised dimension reduction for large-scale \"omics\" data with censored survival outcomes under possible non-proportional hazards 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.