Functional similarity of non-coding regions is revealed in phylogenetic average motif score representations
Alam, A.; Duncan, A. G.; Mitchell, J. A.; Moses, A. M.
Show abstract
Here we frame the cis-regulatory code (that connects the regulatory functions of non-coding regions, such as promoters and UTRs, to their DNA sequences) as a representation building problem. Representation learning has emerged as a new approach to understand function of DNA and proteins, by projecting sequences into high-dimensional feature spaces, where the features are learned from data by a neural network. Inspired by these approaches, we seek to define a feature space where non-coding regions with similar regulatory functions are nearby each other. As a first attempt, we engineered features based on matches to biochemically characterized regulatory motifs in the DNA sequences of non-coding regions. Remarkably, we found that functionally similar promoters and 3 UTRs could be grouped together in a feature space defined by simple averages of the best match scores in (unaligned) orthologous non-coding regions, which we refer to as phylogenetic average motif scores. Perhaps most important, because this feature space is based on known motifs and not fit to any data, it is fully interpretable and not limited to any particular cell type or experimental context. We find that we can read off known regulatory relationships and evolutionary rewiring from visualizations of phylogenetic average motif score representations, and that predicted regulatory interactions based on neighbors in the feature space are borne out in transcription factor deletion experiments. Phylogenetic averages of match scores to known motifs is a baseline for representation learning applied to non-coding sequences, and may continue to improve as databases of motifs become more complete.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Discovering molecular features of intrinsically disordered regions by using evolution for contrastive learning 95%
- Computational identification and experimental characterization of preferred downstream positions in human core promoters 95%
- CoVar: A generalizable machine learning approach to identify the coordinated regulators driving variational gene expression 94%
Similar papers in this journal
- Promoter scanning during transcription initiation in Saccharomyces cerevisiae: Pol II in the "shooting gallery" 96%
- Polygraph: A Software Framework for the Systematic Assessment of Synthetic Regulatory DNA Elements 95%
- Evidence for the role of transcription factors in the co-transcriptional regulation of intron retention 95%
Similar papers in this journal
- ERC 2.0 - evolutionary rate covariation update provides more powerful inference of functional interactions across large phylogenies 94%
- Bayesian Optimized sample-specific Networks Obtained ByOmics data (BONOBO) 93%
- The origin and evolution of a distinct mechanism of transcription initiation in yeasts 93%
Similar papers in this journal
Similar papers in this journal
- Mapping the architecture of regulatory variation provides insights into the evolution of complex traits 95%
- Massively Parallel Analysis of Human 3' UTRs Reveals that AU-Rich Element Length and Registration Predict mRNA Destabilization 93%
- Restricted maximum-likelihood method for learning latent variance components in gene expression data with known and unknown confounders 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.