MotiMul: A significant discriminative sequence motif discovery algorithm with multiple testing correction
Mori, K.; Ozaki, H.; Fukunaga, T.
Show abstract
Sequence motifs play essential roles in intermolecular interactions such as DNA-protein interactions. The discovery of novel sequence motifs is therefore crucial for revealing gene functions. Various bioinformatics tools have been developed for finding sequence motifs, but until now there has been no software based on statistical hypothesis testing with statistically sound multiple testing correction. Existing software therefore could not control for the type-1 error rates. This is because, in the sequence motif discovery problem, conventional multiple testing correction methods produce very low statistical power due to overly-strict correction. We developed MotiMul, which comprehensively finds significant sequence motifs using statistically sound multiple testing correction. Our key idea is the application of Tarones correction, which improves the statistical power of the hypothesis test by ignoring hypotheses that never become statistically significant. For the efficient enumeration of the significant sequence motifs, we integrated a variant of the PrefixSpan algorithm with Tarones correction. Simulation and empirical dataset analysis showed that MotiMul is a powerful method for finding biologically meaningful sequence motifs. The source code of MotiMul is freely available at https://github.com/ko-ichimo-ri/MotiMul.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- GradHC: Highly Reliable Gradual Hash-based Clustering for DNA Storage Systems 95%
- A Pseudo-Temporal Causality Approach to Identifying miRNA-mRNA Interactions During Biological Processes 95%
- Differential co-expression network analysis with DCoNA reveals isomiR targeting aberrations in prostate cancer 95%
Similar papers in this journal
- ProALIGN: Directly learning alignments for protein structure prediction via exploiting context-specific alignment motifs 94%
- Sensitivity analysis of genome-scale metabolic flux prediction 94%
- Combined topological data analysis and geometric deep learning reveal niches by the quantification of protein binding pockets 93%
Similar papers in this journal
- GCNCDA: A New Method for Predicting CircRNA-Disease Associations Based on Graph Convolutional Network Algorithm 95%
- scPADGRN: A preconditioned ADMM approach for reconstructing dynamic gene regulatory network using single-cell RNA sequencing data 95%
- LMSM: a modular approach for identifying lncRNA related miRNA sponge modules in breast cancer 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.