A feature ranking algorithm for clustering medical data
Shpigelman, E.; Shamir, R.
Show abstract
ObjectiveClustering methods are often applied to electronic medical records (EMR) for various objectives, including the discovery of previously unrecognized disease subtypes. The abundance and redundancy of information in EMR data raises the need to rank the features by their relevance to clustering. MethodsHere we propose FRIGATE, an ensemble feature ranking algorithm for clustering. FRIGATE ranks the features by solving multiple clustering problems on subgroups of features, using game-theoretic principles to rank and weigh features. In every such problem, a Shapley-like framework is utilized to rank a selected set of features. In another version of the algorithm, multiplicative weights are employed to reduce the randomness in feature set selection. The code for the algorithms is available in: https://github.com/Shamir-Lab/FRIGATE. ResultsOn simulated data and on eleven real genomics and EMR datasets, FRIGATE outperforms extant ensemble ranking algorithms, in solution quality and in speed. ConclusionFrigate can improve disease understanding by enabling better subtype discovery from EMR data.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Intrinsic-Dimension analysis for guiding dimensionality reduction and data fusion in multi-omics data processing 95%
- Graph Neural Network Modelling as a potentially effective Method for predicting and analyzing Procedures based on Patient Diagnoses 95%
- Stability of feature selection utilizing Graph Convolutional Neural Network and Layer-wise Relevance Propagation 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.