Mutual Information Reveals a 6 Base Pair Signal in Noncoding DNA
Dannemiller, J. L.
Show abstract
MotivationStatistical dependencies between nucleotides at different positions within a DNA sequence have been used for several purposes including distinguishing coding from noncoding regions of a genome. Coding sequences show correlations within and between codon positions. This study asked whether such correlations between positions separated by short distances might also exist in noncoding DNA. To this end, positional nucleotide dependencies were examined in the promoter regions of four eukaryotic species: Homo sapiens (Hs), Mus musculus (Mm), Drosophila melanogaster (Dm), and Saccharomyces cerevisiae (Sc). The degree of dependency between pairwise positions across a set of aligned sequences was quantified by Mutual Information (MI) and visualized using a novel heatmap method. ResultsMI in promoter sequences aligned at their putative Transcription Start Site (TSS) generally decreased with increasing distance between two positions, but also showed a prominent increase at a distance of 6 base pairs (bp) (i.e., between nucleotides at x and x + 6 in a sequence) in the three multicellular species, but much less so in Sc. This dependency at a distance of 6 bp appears to reflect an N1 ... N7 homonucleotide bias in promoters. AvailabilityR code and data files available at github/dannemil/promoters. Contactdannemil@rice.edu Supplementary informationDannemiller-supplementary-documents.zip
Matching journals
The top 11 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Machine learning techniques for classifying the mutagenic origins of point mutations 92%
- Two promoters integrate multiple enhancer inputs to drive wild-type knirps expression in the D. melanogaster embryo 92%
- Natural variation in codon bias and mRNA folding strength interact synergistically to modify protein expression in Saccharomyces cerevisiae 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.