CausalDRIFT: Causal Dimensionality Reduction via Inference of Feature Treatments for Robust Healthcare Machine Learning
Hasan, K. S.; Mou, J. A.
Show abstract
High-dimensional medical datasets present challenges in feature selection, where traditional methods often prioritize spurious correlations over causally relevant variables, compromising model interpretability and clinical utility. We introduce CausalDRIFT, a causal feature selection algorithm grounded in the Frisch-Waugh-Lovell theorem and Double Machine Learning, which estimates the Average Treatment Effect (ATE) of each feature on clinical outcomes while adjusting for confounders. We evaluated CausalDRIFT against seven baseline methods (PCA, ICA, Elastic Net, RFE, etc.) across four datasets (Heart Disease, Diabetes, Breast Cancer, and PCOS) using XGBoost classifiers, with performance metrics including accuracy, precision, recall, and F1-score. CausalDRIFT achieved competitive performance, notably excelling in datasets with strong causal structure (e.g., 90% accuracy and F1-score of 0.90 on PCOS, outperforming most other methods). It demonstrated superior consistency (lowest standard deviation: 1.19 in Breast Cancer, lowest recall spread in Heart Disease) and robustness to confounding, though it traded marginal predictive gains for interpretability in correlation-dominated datasets. CausalDRIFT excels in high-dimensional, low-sample size (HDLSS) settings, such as the Breast Cancer dataset (569x32), where it achieves 93.9% accuracy and 91.8% F1-score. Statistical analysis (ANOVA, Tukey HSD) confirmed its recall performance was non-inferior to top methods (all p > 0.15), while unsupervised techniques like ICA significantly underperformed (p < 0.05). CausalDRIFT bridges the gap between causal inference and scalable feature selection, offering clinically interpretable and generalizable models. Its ability to prioritize causally actionable features is critical for high-stakes decision-making, and it makes it a promising tool for healthcare AI, particularly in settings like PCOS where etiology is complex.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Using computable knowledge mined from the literature to elucidate confounders for EHR-based pharmacovigilance 95%
- Mining for Equitable Health: Assessing the Impact of Missing Data in Electronic Health Records 95%
- A methodology of phenotyping ICU patients from EHR data: high-fidelity, personalized, and interpretable phenotypes estimation 95%
Similar papers in this journal
- Addressing Label Noise for Electronic Health Records: Insights from Computer Vision for Tabular Data 95%
- Causal Analysis for Multivariate Integrated Clinical and Environmental Exposures Data 95%
- OASIS+: leveraging machine learning to improve the prognostic accuracy of OASIS severity score for predicting in-hospital mortality 94%
Similar papers in this journal
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 95%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 94%
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 94%
Similar papers in this journal
- Modeling physician variability to prioritize relevant medical record information 96%
- Using indication embeddings to represent patient health for drug safety studies 94%
- Characterizing subgroup performance of probabilistic phenotype algorithms within older adults: A case study for dementia, mild cognitive impairment, and Alzheimer’s and Parkinson’s diseases 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.