MINDFUL: A Method to Identify Novel and Diverse Signals with Fast, Unsupervised Learning
Parulekar, M.; Narlikar, L.
Show abstract
With rapid advances in experimental methods that map transcription start sites (TSSs) at a high resolution, there is a need to characterize the sequence diversity of TSS neighborhoods. Most current techniques scan for previously discovered elements, such as the TATA box, the INR motif, CpG islands, etc. to categorize promoters into different classes. Reliance on such elements hinders the discovery of novel elements. On the other hand, methods that use standard motif discovery to discover de novo promoter elements are also limited by the fact that a motif is picked up only if it is over-represented in the dataset. An element that appears only in a small set of promoters can thus be missed. We previously developed a clustering-based approach that uses no prior knowledge of elements to solve this problem [1]. That method uses Gibbs sampling to learn the model parameters, but is untenable on large datasets. Here we propose a new, fast method called MINDFUL, that uses a greedy k-means-like approach to cluster promoters aligned by TSSs into diverse classes, while also learning the optimal value of k. It is general enough to be used for any data that has categorical variables, and is not restricted to DNA.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Potpourri: An Epistasis Test Prioritization Algorithm via Diverse SNP Selection 93%
- An Efficient, Scalable and Exact Representation of High-Dimensional Color Information Enabled via de Bruijn Graph Search 93%
- Enabling inference for context-dependent models of mutation by bounding the propagation of dependency 93%
Similar papers in this journal
- Inferring Tumor Progression in Large Datasets 95%
- Reconstruction Set Test (RESET): a computationally efficient method for single sample gene set testing based on randomized reduced rank reconstruction error 94%
- Learning massive interpretable gene regulatory networks of the human brain by merging Bayesian Networks 94%
Similar papers in this journal
- GeneSNAKE: a Python package for benchmarking and simulation of gene regulatory networks and perturbation-induced expression data 95%
- Batch-effect correction in single-cell RNA sequencing data using JIVE 95%
- UTRGAN: Learning to Generate 5' UTR Sequences for Optimized Translation Efficiency and Gene Expression 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.