GRID-seq assisted computational prediction of transcription factor binding motifs
Veldsman, W. P.
Show abstract
Experimental validation of computationally predicted transcription factor binding motifs is desirable. Increased RNA levels in the vicinity of predicted protein-chromosomal binding motifs intuitively suggest regulatory activity. With this intuition in mind, the approach presented here juxtaposes publicly available experimentally derived GRID-seq data with binding motif predictions computationally determined by deep learning models. The aim is to demonstrate the feasibility of using RNA-sequencing data to improve binding motif prediction accuracy. Publicly available GRID-seq scores and computed DeepBind scores could be aggregated by chromosomal region and anomalies within the aggregated data could be detected using mahalanobis distance analysis. A mantels test of matrices containing pairwise hamming distances showed significant differences between 1) randomly ranked sequences, 2) sequences ranked by non-GRID-seq assisted scores, and 3) sequences ranked by GRID-seq assisted scores. Plots of mahalanobis ranked binding motifs revealed an inversely proportional relationship between GRID-seq scores and DeepBind scores. Data points with high DeepBind scores but low GRID-seq scores had no DNAse hypersensitivity clusters annotated to their respective sequences. However, DNase hypersensitivity was observed for high scoring DeepBind motifs with moderate GRID-seq scores. Binding motifs of interest were recognized by their deviance from the inversely proportional tendency, and the underlying context sequences of these predicted motifs were on occasion associated with DNAse hypersensitivity unlike the most highly ranked motif scores when DeepBind was used in isolation. This article presents a novel combinatory approach to predict functional protein-chromosomal binding motifs. The two underlying methods are based on recent developments in the fields of RNA sequencing and deep learning, respectively. They are shown to be suited for synergistic use, which broadens the scope of their respective applications.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- tRForest: a novel random forest-based algorithm for tRNA-derived fragment target prediction 94%
- Genomic background sequences systematically outperform synthetic ones in de novo motif discovery for ChIP-seq data 94%
- FLYNC: A Machine Learning-Driven Framework for Discovering Long Non-Coding RNAs in Drosophila melanogaster 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.