Using bootstrap procedures for testing the modular partition inferred via leading eigenvector community detection algorithm
Dzeverin, I.; Vertsimakha, O.
Show abstract
Modularity and modular structures can be recognized at various levels of biological organization and in various domains of studies. Recently, algorithms based on network analysis came into focus. And while such a framework is a powerful tool in studying modular structure, those methods usually pose a problem of assessing statistical support for the obtained modular structures. One of the widely applied methods is the leading eigenvector, or Newmans spectral community detection algorithm. We conduct a brief overview of the method, including a comparison with some other community detection algorithms and explore a possible fine-tuning procedure. Finally, we propose an adapted bootstrap-based procedure based on Shimodairas multiscale bootstrap algorithm to derive approximately unbiased p-values for the module partitions of observations datasets. The proposed procedure also gives a lot of freedom to the researcher in constructing the network construction from the raw numeric data, and can be applied to various types of data and used in diverse problems concerning modular structure. We provide an R language code for all the calculations and the visualization of the obtained results for the researchers interested in using the procedure.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- DnoisE: Distance denoising by Entropy. An open-source parallelizable alternative for denoising sequence datasets 90%
- Linking Parasitism to Network Centrality and the Impact of Sampling Bias in its Interpretation 90%
- Evaluating probabilistic programming and fast variational Bayesian inference in phylogenetics 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.