A Deep Neural Network Two-part Model and Feature Importance Test for Semi-continuous Data
Zou, B.; Mi, X.; Xenakis, J.; Wu, D.; Hu, J.; Zou, F.
Show abstract
Semi-continuous data frequently arise in clinical practice. For example, while many surgical patients suffer from varying degrees of acute postoperative pain (POP) post surgery (i.e., POP score > 0), others experience none (i.e., POP score = 0), indicating the existence of two distinct data processes at play. Existing parametric or semi-parametric two-part modeling methods for this type of semicontinuous data can fail to appropriately model these two underlying data processes as such methods rely heavily on (generalized) linear additive assumptions. However, many factors may interact to jointly influence the experience of POP non-additively and non-linearly. Motivated by this challenge and inspired by the flexibility of deep neural networks (DNN) to accurately approximate complex functions universally, we derive a DNN-based two-part model by adapting the conventional DNN methods by adding two additional components: a bootstrapping procedure along with a filtering algorithm to boost the stability of the conventional DNN, an approach we denote as sDNN. To improve the interpretability and transparency of sDNN, we further derive a feature importance testing procedure to identify important features contributing to the outcome measurements of the two data processes, denoting this approach fsDNN. We show that fsDNN not only offers a valid feature importance test but also that using the identified features can further improve the predictive performance of sDNN. The proposed sDNN- and fsDNN-based twopart models are applied to the analysis of real data from a POP study, in which application they clearly demonstrate advantages over the existing parametric and semi-parametric two-part models. Further, we conduct extensive numerical studies to demonstrate that sDNN and fsDNN consistently outperform the existing two-part models regardless of the data complexity. An R package implementing the proposed methods has been developed and deposited on GitHub (https://github.com/SkadiEye/fsDNN).
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- High-dimensional Biomarker Identification for Scalable and Interpretable Disease Prediction via Machine Learning Models 96%
- Incorporating Prior Knowledge into Regularized Regression 96%
- Joint eQTL mapping and Inference of Gene Regulatory Network Improves Power of Detecting both cis- and trans-eQTLs 95%
Similar papers in this journal
- Using random forests to uncover the predictive power of distance-varying cell interactions in tumor microenvironments 95%
- Model guided trait-specific co-expression network estimation as a new perspective for identifying molecular interactions and pathways 95%
- Inferring Tumor Progression in Large Datasets 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- A Regularized Cox Hierarchical Model for Incorporating Annotation Information in Predictive Omic Studies 95%
- Estimating Microbial Interaction Network:Zero-inflated Latent Ising Model Based Approach 94%
- A compact encoding of the genome suitable for machine learning prediction of traits and genetic risk scores. 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.