KDM: embedding DNA/RNA motifs and sequences in a shared k-mer space for unified discovery, analysis and binding prediction
Fumagalli, L.; Becchi, T.; Cereda, M.; Pozzoli, U.
Show abstract
Motif discovery and binding-site prediction in DNA and RNA sequences are central tasks in regulatory genomics, yet the methodological landscape is split between interpretable but rigid position weight matrices (PWMs) and high-performing but opaque machine-learning models. We present KDM, a unifying framework in which both motifs and sequences are represented as probability distributions over a shared k-mer dictionary, embedded via the Hellinger transformation. This common geometry enables motif-sequence scoring, motif-motif comparison, de novo discovery, and binding prediction with a single primitive, the Bhattacharyya coefficient. We instantiate four tools on this representation: KDMMap for positional enrichment analysis, KDMMatch for information-content-aware motif matching, KDMFind for unsupervised motif discovery via projective non-negative matrix factorization, and KDM-LRLM for binding prediction with Lasso-regularized logistic regression. Across 1,324 transcription-factor ChIP-seq and 161 RBP eCLIP experiments, KDMMap matches CentriMos motif rankings in 84% of TF and 79% of RBP experiments, and KDMMatch agrees with Tomtom on motif annotation in 74.5% of TFs. On binding prediction across four datasets covering 2,475 experiments, KDM-LRLM matches or exceeds eight deep-learning and three k-mer-based competitors. Notably, AI methods overtake k-mer methods only in the top quartile of training-set size, indicating that data scale, not architecture, drives the recent dominance of deep models. KDM provides a single interpretable representation across the full motif-analysis workflow.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Differential Analysis of RNA Structure Probing Experiments at Nucleotide Resolution: Uncovering Regulatory Functions of RNA Structure 97%
- Massively parallel reporter perturbation assay uncovers temporal regulatory architecture during neural differentiation 96%
- PACS allows comprehensive dissection of multiple factors governing chromatin accessibility from snATAC-seq data 96%
Similar papers in this journal
- Evaluating the representational power of pre-trained DNA language models for regulatory genomics 97%
- An interpretable bimodal neural network characterizes the sequence and preexisting chromatin predictors of induced TF binding 97%
- Correcting gradient-based interpretations of deep neural networks for genomics 96%
Similar papers in this journal
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 96%
- DeepCLIP: Predicting the effect of mutations on protein-RNA binding with Deep Learning 96%
- Deciphering the 3D genome organization across species from Hi-C data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.