Hierarchical Clustering with an Ensemble of Principle Component Trees for Interpretable Patient Stratification
Pfeifer, B.; Bloice, M. D.; Sirocchi, C.; Loecher, M.
Show abstract
Patient stratification plays a crucial role in personalized medicine by identifying distinct subgroups of patients based on their molecular and/or clinical characteristics. However, many unsupervised machine learning-based stratification techniques fail to identify the essential biomarker traits associated with each patient group. In this paper, we present a novel approach for interpretable patient stratification using hierarchical ensemble clustering. Our method leverages feature sampling in conjunction with principal component analysis (PCA) to capture the most significant patterns and contributing biomarkers. We demonstrate the effectiveness of our approach using machine learning benchmark datasets and real-world data from The Cancer Genome Atlas (TCGA), showcasing the improved interpretability of the detected patient clusters.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Stability of feature selection utilizing Graph Convolutional Neural Network and Layer-wise Relevance Propagation 95%
- Graph Neural Network Modelling as a potentially effective Method for predicting and analyzing Procedures based on Patient Diagnoses 95%
- Intrinsic-Dimension analysis for guiding dimensionality reduction and data fusion in multi-omics data processing 94%
Similar papers in this journal
- A methodology of phenotyping ICU patients from EHR data: high-fidelity, personalized, and interpretable phenotypes estimation 93%
- Individual Reference Intervals for Personalized Interpretation of Clinical and Metabolomics Measurements 92%
- A Utility-Based Machine Learning-Driven Personalized Lifestyle Recommendation for Cardiovascular Disease Prevention 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.