Hobotnica: exploring molecular signature quality
Stupnikov, A.; Sizykh, A.; Favorov, A.; Afsari, B.; Wheelan, S. J.; Marchionni, L.; Medvedeva, Y. A.
Show abstract
A Molecular Features Set (MFS), is a result of vast diversity of bioinformatics pipelines. In case when MFS is used for further analysis to distinguish between phenotypes, it is often referred to as a signature. Lack of the "gold standard" for most experimental data modalities makes it hard to provide valid estimation for a particular MFSs quality. Yet, this goal can partially be achieved by analyzing inner-sample Distance Matrix (DM) and their power to distinguish between phenotypes. The quality of a DM can be assessed by summarizing its power to quantify the differences of inner-phenotype and outer-phenotype distances. This estimation of the DM quality can be construed as a measure of the MFSs quality. Here we propose Hobotnica, an approach to estimate MFSs quality by their ability to stratify data, and assign them significance scores, that allows for collating various signatures and comparing their quality for contrasting groups.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Benchmarking feature selection and feature extraction methods to improve the performances of machine-learning algorithms for patient classification using metabolomics biomedical data. 94%
- Thresholding Gini Variable Importance with a single trained Random Forest: An Empirical Bayes Approach 94%
- iMDA-BN: Identification of miRNA-Disease Associations based on the Biological Network and Graph Embedding Algorithm 94%
Similar papers in this journal
- Model guided trait-specific co-expression network estimation as a new perspective for identifying molecular interactions and pathways 95%
- Methodological Challenges in Translational Drug Response Modeling in Cancer 95%
- DGCyTOF: deep learning with graphic cluster visualization to predict cell types of single cell mass cytometry data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.