Back

A feature ranking algorithm for clustering medical data

Shpigelman, E.; Shamir, R.

2023-09-30 health informatics
10.1101/2023.09.30.23296349 medRxiv
Show abstract

ObjectiveClustering methods are often applied to electronic medical records (EMR) for various objectives, including the discovery of previously unrecognized disease subtypes. The abundance and redundancy of information in EMR data raises the need to rank the features by their relevance to clustering. MethodsHere we propose FRIGATE, an ensemble feature ranking algorithm for clustering. FRIGATE ranks the features by solving multiple clustering problems on subgroups of features, using game-theoretic principles to rank and weigh features. In every such problem, a Shapley-like framework is utilized to rank a selected set of features. In another version of the algorithm, multiplicative weights are employed to reduce the randomness in feature set selection. The code for the algorithms is available in: https://github.com/Shamir-Lab/FRIGATE. ResultsOn simulated data and on eleven real genomics and EMR datasets, FRIGATE outperforms extant ensemble ranking algorithms, in solution quality and in speed. ConclusionFrigate can improve disease understanding by enabling better subtype discovery from EMR data.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.