Consensus Clustering for Robust Bioinformatics Analysis
Yousefi, B.; Schwikowski, B.
Show abstract
Clustering plays an important role in a multitude of bioinformatics applications, including protein function prediction, population genetics, and gene expression analysis. The results of most clustering algorithms are sensitive to variations of the input data, the clustering algorithm and its parameters, and individual datasets. Consensus clustering (CC) is an extension to clustering algorithms that aims to construct a robust result from those clustering features that are invariant under the above sources of variation. As part of CC, stability scores can provide an idea of the degree of reliability of the resulting clustering. This review structures the CC approaches in the literature into three principal types, introduces and illustrates the concept of stability scores, and illustrates the use of CC in applications to simulated and real-world gene expression datasets. Open-source R implementations for each of these CC algorithms are available in the GitHub repository: https://github.com/behnam-yousefi/ConsensusClustering
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Gene regulation network inference using k-nearest neighbor-based mutual information estimation- Revisiting an old DREAM 95%
- DeepMF: Deciphering the Latent Patterns in Omics Profiles with a Deep Learning Method 95%
- eSVD-DE: Cohort-wide differential expression in single-cell RNA-seq data using exponential-family embeddings 94%
Similar papers in this journal
- SillyPutty: Improved clustering by optimizing the silhouette width 96%
- Compressive Big Data Analytics: An Ensemble Meta-Algorithm for High-dimensional Multisource Datasets 95%
- Voting-based integration algorithm improves causal network learning from interventional and observational data: an application to cell signaling network inference 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.