Direct causal variable discovery leveraging the invariance principle: application in biomedical studies
Yin, L.; Liu, M.; Shi, Y.; Qiu, J.; So, H.-C.
Show abstract
Accurate identification of direct causal (parental) variables for a target is of primary interest in many applications, especially in biomedical sciences. It could promote our understanding of disease mechanisms, and facilitate the discovery of new biomarkers and therapeutic targets for clinical traits. However, standard machine learning approaches often identify spurious associations, while existing causal inference methods for direct causal variables can be computationally infeasible for high-dimensional biomedical data. Here, we proposed a novel and efficient two-stage approach (I-GCM) to discover direct causal variables (including genetic and clinical variables) for clinical outcomes. The method first employs the PC-simple algorithm for feature screening, then leverages the principle of causal invariance across different environments. Causal relationships are robustly identified by testing for changes in the generalized covariance measure (GCM), calculated using flexible gradient-boosted tree models. We first verified the proposed method through extensive simulations. I-GCM constantly yielded high precision (positive predictive value) and specificity while maintaining satisfactory sensitivity in general, and consistently outperformed a standard method. Notably, the precision was larger than 90% in our simulated scenarios, even in high-dimensional settings. We then applied I-GCM to the UK-Biobank, analyzing genetic and clinical data to identify causal factors for COVID-19 infection/severity and lipid traits (HDL, Triglycerides). The analysis successfully recovered many known clinical risk factors, validating the methods real-world performance, and uncovered novel putative causal genes and biological pathways supported by existing literature. Importantly, our work pioneers the application of the invariance principle for causal inference in biomedical/clinical studies, and suggests a new avenue for causal discovery in these settings.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Finding disease modules for cancer and COVID-19 in gene co-expression networks with the Core&Peel method 95%
- DeepInsight-3D for precision oncology: an improved anti-cancer drug response prediction from high-dimensional multi-omics data with convolutional neural networks 95%
- Stochastic LASSO for extremely high-dimensional genomic data 94%
Similar papers in this journal
- A Generalized Higher-order Correlation Analysis Framework for Multi-Omics Network Inference 96%
- MENDELSEEK: An algorithm that predicts Mendelian Genes and elucidates what makes them special 95%
- Explainable deep transfer learning model for disease risk prediction using high-dimensional genomic data 95%
Similar papers in this journal
- Disease Network Delineates the Disease Progression Profile of Cardiovascular Diseases 93%
- Automated Interpretable Discovery of Heterogeneous Treatment Effectiveness: A Covid-19 Case Study 93%
- Individual Reference Intervals for Personalized Interpretation of Clinical and Metabolomics Measurements 93%
Similar papers in this journal
- Penalized reduced rank regression for multi-outcome survival data supports a common metabolic risk score for age-related diseases 94%
- Efficient Estimation of Indirect Effects in Case-Control Studies Using a Unified Likelihood Framework 94%
- Using a supervised principal components analysis for variable selection in high-dimensional datasets reduces false discovery rates 93%
Similar papers in this journal
- KGRACDA: A Model Based on Knowledge Graph from Recursion and Attention Aggregation for CircRNA-disease Association Prediction 94%
- CPGL: Prediction of compound-protein interaction by integrating graph attention network with long short-term memory neural network 93%
- Genetic analysis of coronary artery disease using tree-based automated machine learning informed by biology-based feature selection 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.