Gene Set Overlap: An Impediment to Achieving High Specificity in Over-representation Analysis
Maleki, F.; Kusalik, A. J.
Show abstract
Gene set analysis methods are widely used to analyze data from high-throughput "omics" technologies. One drawback of these methods is their low specificity or high false positive rate. Over-representation analysis is one of the most commonly used gene set analysis methods. In this paper, we propose a systematic approach to investigate the hypothesis that gene set overlap is an underlying cause of low specificity in over-representation analysis. We quantify gene set overlap and show that it is a ubiquitous phenomenon across gene set databases. Statistical analysis indicates a strong negative correlation between gene set overlap and the specificity of over-representation analysis. We conclude that gene set overlap is an underlying cause of the low specificity. This result highlights the importance of considering gene set overlap in gene set analysis and explains the lack of specificity of methods that ignore gene set overlap. This research also establishes the direction for developing new gene set analysis methods.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- SC-JNMF: Single-cell clustering integrating multiple quantification methods based on joint non-negative matrix factorization 95%
- Detection of spreader nodes and ranking of interacting edges in Human-SARS-CoV protein interaction network 92%
- Network based multifactorial modelling of miRNA-target interactions 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.