KMAP: Kmer Manifold Approximation and Projection for visualizing DNA sequences
Fu, C.; Niskanen, E. A.; Wei, G.; Yang, Z.; Sanvicente-Garcia, M.; Güell, M.; Cheng, L.
Show abstract
Identifying and illustrating patterns in DNA sequences is a crucial task in various biological data analyses. In this task, patterns are often represented by sets of kmers, the fundamental building blocks of DNA sequences. To visually unveil these patterns, we could project each kmer onto a point in two-dimensional (2D) space. However, this projection poses challenges due to the high-dimensional nature of kmers and their unique mathematical properties. Here, we established a mathematical system to address the peculiarities of the kmer manifold. Leveraging this kmer manifold theory, we developed a statistical method named KMAP for detecting kmer patterns and visualizing them in 2D space. We applied KMAP to three distinct datasets to showcase its utility. KMAP achieved a comparable performance to the classical method MEME, with approximately 90% similarity in motif discovery from HT-SELEX data. In the analysis of H3K27ac ChIP-seq data from Ewing Sarcoma (EWS), we found that BACH1, OTX2 and ERG1 might affect EWS prognosis by binding to promoter and enhancer regions across the genome. We also found that FLI1 bound to the enhancer regions after ETV6 degradation, which showed the competitive binding between ETV6 and FLI1. Moreover, KMAP identified four prevalent patterns in gene editing data of the AAVS1 locus, aligning with findings reported in the literature. These applications underscore that KMAP could be a valuable tool across various biological contexts. KMAP is freely available at: https://github.com/chengl7-lab/kmap.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- EpiSAFARI: Sensitive detection of valleys in epigenetic signals for enhancing annotations of functional elements 95%
- simCAS: an embedding-based method for simulating single-cell chromatin accessibility sequencing data 95%
- Clustering single-cell RNA-seq data by rank constrained similarity learning 95%
Similar papers in this journal
- Assessing base-resolution DNA mechanics on the genome scale 96%
- Systematic Prediction of Regulatory Motifs from Human ChIP-Sequencing Data Based on a Deep Learning Framework 95%
- ANANSE: An enhancer network-based computational approach for predicting key transcription factors in cell fate determination 95%
Similar papers in this journal
Similar papers in this journal
- scTensor detects many-to-many cell-cell interactions from single cell RNA-sequencing data 95%
- Identification and Utilization of Copy Number Information for Correcting Hi-C Contact Map of Cancer Cell Line 94%
- Gene2role: a role-based gene embedding method for comparative analysis of signed gene regulatory networks 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.