NIMAA: an R/CRAN package to accomplish NomInal data Mining AnAlysis
Jafari, M.; Chen, C.; Mirzaie, M.; Tang, J.
Show abstract
SummaryNominal data is data that has been "labeled" and can be designated into a number of non-overlapping unordered groups. The analysis of this type of data is often trivial because it is not feasible to conduct extensive numerical methods on this type of data. Graphs or networks, on the other hand, are comprised of sets of nodes and edges that can also be considered as nominal variables. By integrating graph theory and data mining approaches, we offer the R package NIMAA to define a nominal data-mining pipeline to explore more information. Using nominal variables in a dataset, NIMAA provides functions for constructing weighted and unweighted bipartite graphs, analysing the similarity of labels in nominal variables, clustering labels or categories to super-labels, validating clustering results, predicting bipartite edges by missing weight imputation, and providing a variety of visualization tools. Here, we also indicated the application of nominal data mining in a biological dataset with well-riched nominal variables. AvailabilityNIMAAs official release and the beta update are available on CRAN and Github, respectively. URLs: https://CRAN.R-project.org/package=NIMAA and https://github.com/jafarilab/NIMAA Contactmohieddin.jafari@helsinki.fi; jing.tang@helisnki.fi ContributionsMJ conceived the study and developed the models, MJ and CC adopted and implemented the methods, MM improved the methods, JT provided the funding, MJ, CC, MM and JT wrote the paper.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- eDAVE - extension of GDC Data Analysis, Visualization, and Exploration Tools 92%
- Thresholding Gini Variable Importance with a single trained Random Forest: An Empirical Bayes Approach 92%
- Benchmarking feature selection and feature extraction methods to improve the performances of machine-learning algorithms for patient classification using metabolomics biomedical data. 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.